Using Tokmine with an account
An account turns a throwaway session into a persistent, balanced, billable setup: stable URLs, several machines behind one URL, and per-person tokens.
1. Get an agent token
Sign up, then open the portal and generate an agent token on the machines page. It is shown once. It identifies your workspace to Tokmine; it is not what API callers use.
2. Log in on each machine
tokmine login --token <agent-token>
tokmine status # confirms: authenticated (token configured)
tokminetokmine login stores the token in the config file (see CLI reference) with owner-only permissions. Alternatives: the TOKMINE_TOKEN environment variable (handy for containers and services), or tokmine login with no arguments to paste it interactively. Undo with tokmine logout; the machine id is kept.
From then on tokmine connects as a registered machine with a persistent public URL and no 30-minute limit. Use --name to give the machine a recognisable name; it defaults to the hostname.
Machines
Every machine running the CLI shows up in the portal with its models, CPU, RAM and GPU load, and connection state. Machines are a billing unit. If you hit your plan limit the CLI stops with a clear message instead of retrying forever.
Endpoints
Your workspace always has a default endpoint. On a paid plan you can also create static endpoints with names and URLs you choose, and point machines at them with tokmine --endpoint <slug>. Run tokmine endpoints to list them. Details in Static endpoints and load balancing.
Roles
Every member of the workspace can see endpoints and call them with their own token. Creating, disabling and deleting endpoints, and managing machines and billing, is limited to workspace admins.
Seats and consumer tokens
People and services that call your URL each use their own token, created in the portal. Each token belongs to a seat, so usage and access are per person and can be revoked independently. Send it as Authorization: Bearer, x-api-key, x-goog-api-key or x-token, so stock SDKs work unchanged.
Load balancing across machines
Point several machines at the same workspace and they share one public URL. Each request is routed to a healthy machine that serves the requested model, favouring the least loaded one using the CPU, RAM, GPU and in-flight request counts the CLI reports continuously. If a machine drops, traffic shifts to the others; when it returns it rejoins automatically.
# gpu-box-1
tokmine --name gpu-box-1
# gpu-box-2
tokmine --name gpu-box-2
# both now serve the same URL, https://tunnel.tokmine.ai/t/<your-slug>
# with several static endpoints, add --endpoint <slug> to choose oneNext: Static endpoints and the CLI reference.