Decentralized compute for LLMs
A traffic manager for your local AI models
Run tokmine and get a public, authenticated URL for the models on your machines. Traffic is balanced across them by live CPU, RAM and GPU load. OpenAI, Anthropic and Gemini compatible. No port-forwarding.
curl -fsSL https://tokmine.ai/install.sh | shNo account needed for the 30-minute session.
- Ollama
- LM Studio
- vLLM
- llama.cpp
$ tokmine Scanning local model servers... ✓ Ollama :11434 openai, ollama 3 models ✓ LM Studio :1234 openai 1 model ✓ Tunnel connected (wss) Public URL https://tunnel.tokmine.ai/t/k7x2-m9qa Token tm_anon_•••••••••••• Free session expires in 30:00. tokmine login to keep it. Chat https://app.tokmine.ai/chat#url=…
Works with the runners and SDKs you already use
- Ollama
- LM Studio
- vLLM
- llama.cpp
- OpenAI API
- Anthropic API
- Gemini API
Install, run, use the URL
Install the CLI
One copy-paste on macOS, Linux or Windows. A single static binary, no runtime needed.
Run tokmine
It discovers your local model servers, detects which API dialect each one speaks, and opens an outbound tunnel. No model server yet? tokmine install ollama sets one up. Nothing to configure, no inbound ports.
Use the URL
Point any OpenAI, Anthropic or Gemini SDK at https://tunnel.tokmine.ai/t/<slug> with your token. Or open the chat link it prints to try the model in your browser. Works with Claude Code, Codex, OpenCode, Aider and more. On a paid plan, create a static URL like /t/acme-prod that several machines can share.
More than a tunnel
Tokmine knows what each machine is doing, so requests go where there is headroom.
Balanced by live load
Each machine reports CPU, RAM, GPU and VRAM use, queue depth and tokens per second. Requests go to the machine best able to serve that model right now.
Routed by model name
Ask for a model and Tokmine finds the machines that serve it, preferring ones that already have it loaded.
Fails over quietly
Unhealthy or draining machines are skipped, and a request that fails before its response starts is retried on another machine.
Every dialect, one URL
OpenAI, Anthropic and Gemini style requests work against the same URL, whichever runner sits behind it.
Your models, reachable everywhere
Share a model with teammates
Give each person their own token instead of a VPN and a wiki page of IP addresses.
Use your home GPU from anywhere
Keep the big model on the desktop at home and call it from the laptop on the train.
Cloud GPU boxes
Run the CLI on any rented VM and add machines to the pool as you need capacity.
CI and agents
Let pipelines and agent frameworks hit local models with a stable, authenticated endpoint.
Public URL, private machine
Authenticated by default
There are no open endpoints. Every request needs a consumer token, and the same tokens work with stock SDKs.
Outbound-only tunnel
The CLI dials out over TLS. Your machine never accepts inbound connections from the internet.
No LAN access
Only the model servers the CLI discovered or you configured are reachable, and only their allow-listed API paths. No arbitrary hosts, no SSRF into your network.
Nothing sensitive in logs
Tokens are stored hashed and shown once. Request bodies and tokens are never logged.