Decentralized compute for LLMs

A traffic manager for your local AI models

Run tokmine and get a public, authenticated URL for the models on your machines. Traffic is balanced across them by live CPU, RAM and GPU load. OpenAI, Anthropic and Gemini compatible. No port-forwarding.

bash
curl -fsSL https://tokmine.ai/install.sh | sh

No account needed for the 30-minute session.

  • Ollama
  • LM Studio
  • vLLM
  • llama.cpp

Works with the runners and SDKs you already use

  • Ollama
  • LM Studio
  • vLLM
  • llama.cpp
  • OpenAI API
  • Anthropic API
  • Gemini API
How it works

Install, run, use the URL

  • Install the CLI

    One copy-paste on macOS, Linux or Windows. A single static binary, no runtime needed.

  • Run tokmine

    It discovers your local model servers, detects which API dialect each one speaks, and opens an outbound tunnel. No model server yet? tokmine install ollama sets one up. Nothing to configure, no inbound ports.

  • Use the URL

    Point any OpenAI, Anthropic or Gemini SDK at https://tunnel.tokmine.ai/t/<slug> with your token. Or open the chat link it prints to try the model in your browser. Works with Claude Code, Codex, OpenCode, Aider and more. On a paid plan, create a static URL like /t/acme-prod that several machines can share.

Traffic management

More than a tunnel

Tokmine knows what each machine is doing, so requests go where there is headroom.

  • Balanced by live load

    Each machine reports CPU, RAM, GPU and VRAM use, queue depth and tokens per second. Requests go to the machine best able to serve that model right now.

  • Routed by model name

    Ask for a model and Tokmine finds the machines that serve it, preferring ones that already have it loaded.

  • Fails over quietly

    Unhealthy or draining machines are skipped, and a request that fails before its response starts is retried on another machine.

  • Every dialect, one URL

    OpenAI, Anthropic and Gemini style requests work against the same URL, whichever runner sits behind it.

Use cases

Your models, reachable everywhere

  • Share a model with teammates

    Give each person their own token instead of a VPN and a wiki page of IP addresses.

  • Use your home GPU from anywhere

    Keep the big model on the desktop at home and call it from the laptop on the train.

  • Cloud GPU boxes

    Run the CLI on any rented VM and add machines to the pool as you need capacity.

  • CI and agents

    Let pipelines and agent frameworks hit local models with a stable, authenticated endpoint.

Security

Public URL, private machine

  • Authenticated by default

    There are no open endpoints. Every request needs a consumer token, and the same tokens work with stock SDKs.

  • Outbound-only tunnel

    The CLI dials out over TLS. Your machine never accepts inbound connections from the internet.

  • No LAN access

    Only the model servers the CLI discovered or you configured are reachable, and only their allow-listed API paths. No arbitrary hosts, no SSRF into your network.

  • Nothing sensitive in logs

    Tokens are stored hashed and shown once. Request bodies and tokens are never logged.