Docs
Tokmine documentation
Tokmine is a traffic manager for local AI models. You run the tokmine CLI on each machine that hosts models, and Tokmine gives you a public, authenticated URL that routes requests to the healthiest machine serving the model.
- Install the CLI on Linux, macOS, Windows or BSD, and see what the installer does.
- Quickstart: a working URL in under a minute, no account.
- Install a local model:
tokmine installsets up Ollama and a model if you have none. - Browser chat: try your endpoint from the link the CLI prints.
- Using an account:
tokmine login, machines, seats and load balancing. - Static endpoints: stable named URLs and load balancing across several machines (paid plan).
- CLI reference: every command, flag, environment variable and the config file.
- Supported runners: how discovery works and how to add remote servers.
- API compatibility: OpenAI, Anthropic and Gemini SDK snippets.
- AI tools and clients: settings for Claude Code, Codex, OpenCode, Aider, Continue, Zed and more.
- Security: what is exposed, and what is not.
- Troubleshooting and FAQ.
Looking for a specific build? See the downloads page.
Concepts
| Machine | A computer running local model servers, on which you run the CLI. Billing unit. |
|---|---|
| Seat | A user who calls the public API through your machines, with their own token. Billing unit. |
| Agent token | Authenticates the CLI to Tokmine (x-token). Identifies you and your workspace. |
| Consumer token | What API callers send to the public URL. Accepted as Authorization: Bearer, x-api-key, x-goog-api-key or x-token, so stock SDKs work. |
| Endpoint | A stable public URL (/t/<slug>) fronting a pool of machines. Every workspace has a default one; paid plans add static, named endpoints. |
| Provider | A local model runner found on a machine (Ollama, LM Studio, vLLM, llama.cpp, ...) and the API dialects it speaks. |
FAQ
Do I need to open ports or configure my router?
No. The CLI opens an outbound TLS connection to Tokmine. Nothing listens on the public internet from your machine.Which model runners are supported?
Anything that exposes an OpenAI, Anthropic or Gemini compatible API, or Ollama's native API. The CLI auto-detects Ollama, LM Studio, vLLM, llama.cpp and other common runners on their default ports, and you can point it at others with --provider.What happens when my free session ends?
After 30 minutes the tunnel closes and the URL is freed. Log in with a token to keep a persistent machine.Troubleshooting: the CLI finds no models.
Make sure your runner is running and listening on localhost. If you have no runner, run tokmine install ollama. Run tokmine models to see what was discovered, or pass --provider http://host:port for a non-default address.Troubleshooting: I get 401 from the public URL.
Send the consumer token printed by the CLI (anonymous) or your own account token (authenticated) using any of the supported headers.