Docs

Quickstart

From zero to a working public URL for your local model, no account needed.

1. Start a model server

Anything works. With Ollama:

bash
ollama serve
ollama pull llama3.2

No local model server yet? Run tokmine install ollama to set one up, see Install a local model.

2. Install and run the CLI

See Install, then run:

bash
tokmine

The CLI scans localhost for model servers, detects which API dialects each one speaks, opens an outbound tunnel and prints what it found:

text
✓ Ollama  :11434  openai, ollama  1 model
✓ Tunnel connected

Public URL  https://tunnel.tokmine.ai/t/k7x2-m9qa
Token       tm_anon_...   (shown once)
Expires     in 30:00
Chat        https://app.tokmine.ai/chat#url=https%3A%2F%2F...&token=...

Anonymous 30-minute session

With no token configured, Tokmine creates a free 30-minute session. The URL and the consumer token are printed when the tunnel comes up, and the countdown shows when it expires. Anonymous sessions are rate limited, and only one is allowed per IP address at a time.

The session is saved locally. If you stop and restart tokmine, or your laptop sleeps, it resumes the same session: same URL, same token, same expiry (the CLI prints resumed previous anonymous session). Once it expires, the next run creates a fresh one.

The last line is a link to a web chat for your endpoint. Open it to try your model in the browser, see Test your endpoint in the browser. The link contains your token, so do not share it. When you are logged in, the link has no token.

3. Call it

Public URLs have the form https://tunnel.tokmine.ai/t/<slug>. Use the URL your CLI printed.

bash
curl https://tunnel.tokmine.ai/t/k7x2-m9qa/v1/chat/completions \
  -H "Authorization: Bearer $TOKMINE_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"model":"llama3.2","messages":[{"role":"user","content":"Hello"}]}'

Or point an official SDK at it. Set the token as the API key:

python
from openai import OpenAI

client = OpenAI(base_url="https://tunnel.tokmine.ai/t/k7x2-m9qa/v1", api_key="<token>")
print(client.chat.completions.create(
    model="llama3.2",
    messages=[{"role": "user", "content": "Hello"}],
).choices[0].message.content)
python
import anthropic

client = anthropic.Anthropic(base_url="https://tunnel.tokmine.ai/t/k7x2-m9qa", api_key="<token>")
print(client.messages.create(
    model="llama3.2", max_tokens=256,
    messages=[{"role": "user", "content": "Hello"}],
).content[0].text)
python
from google import genai
from google.genai import types

client = genai.Client(api_key="<token>", http_options=types.HttpOptions(base_url="https://tunnel.tokmine.ai/t/k7x2-m9qa"))
print(client.models.generate_content(model="llama3.2", contents="Hello").text)

The same URL speaks all three dialects, whichever one your local runner natively supports. Header names, paths and streaming details are in API compatibility.

Using an AI coding tool? See AI tools and clients for Claude Code, Codex, OpenCode, Aider and others.

4. Keep it running: log in

Anonymous sessions end after 30 minutes. For a persistent machine, create an account, generate an agent token in the portal and store it:

bash
tokmine login --token <agent-token>
# or: export TOKMINE_TOKEN=<agent-token>
tokmine

Details in Using Tokmine with an account.

Want a permanent URL?

Paid accounts can create static endpoints: a named URL such as https://tunnel.tokmine.ai/t/acme-prod that stays the same, and that several of your machines can share. Create one on the Endpoints page of the portal, then run tokmine --endpoint acme-prod. Free accounts keep the default endpoint. See Static endpoints and load balancing.

Something not working?

  • No models found: run tokmine models, or add --provider http://host:port. No runner at all? Install a local model.
  • 401 responses: check the token header, see API compatibility.
  • 429 responses: anonymous sessions are rate limited. Log in for higher limits.
  • More in Troubleshooting.