Docs

Use your Tokmine endpoint in AI coding tools and clients

Any tool that lets you change the API base URL can talk to your local models through Tokmine. Every client needs the same three facts: a base URL, a token and a model name.

Find your values

Base URLhttps://tunnel.tokmine.ai/t/<slug>. Clients that speak the OpenAI API want /v1 added: https://tunnel.tokmine.ai/t/<slug>/v1. Anthropic and Gemini clients use the URL without /v1.
TokenAnonymous session: the token the CLI prints. Signed in: your own account token from the portal. Not the agent token.
Model nameExactly as returned by the model list, for example llama3.1:8b.

The URL is printed by tokmine when it connects, and shown with copy buttons on the Endpoints page of the portal. For editor and tool configs use a static endpoint (https://tunnel.tokmine.ai/t/<slug> that never changes), because anonymous session URLs expire after 30 minutes.

Verify before configuring a tool

bash
export TOKMINE_TOKEN=<your token>
curl https://tunnel.tokmine.ai/t/<slug>/v1/models \
  -H "Authorization: Bearer $TOKMINE_TOKEN"

The response lists the models your online machines serve. If that works, any failure in a tool is a client setting. Use these values in the sections below (replace <slug>).

Never commit tokens. Put them in environment variables, your OS keychain or the tool's secret storage, and keep config files that contain a token out of version control. The blocks below read the token from TOKMINE_TOKEN wherever the tool allows it.

On this page

Claude Code

Claude Code speaks the Anthropic Messages API. Point it at your URL without/v1; it appends /v1/messages itself. Anthropic's gateway documentation lists ANTHROPIC_AUTH_TOKEN (sent as Authorization: Bearer) and ANTHROPIC_API_KEY (sent as x-api-key); Tokmine accepts either header.

bash
export ANTHROPIC_BASE_URL=https://tunnel.tokmine.ai/t/<slug>
export ANTHROPIC_AUTH_TOKEN=$TOKMINE_TOKEN
export ANTHROPIC_MODEL=llama3.1:8b
export ANTHROPIC_DEFAULT_HAIKU_MODEL=llama3.1:8b
claude

Or persist it in the env block of ~/.claude/settings.json (%USERPROFILE%\.claude\settings.json on Windows):

json
{
  "env": {
    "ANTHROPIC_BASE_URL": "https://tunnel.tokmine.ai/t/<slug>",
    "ANTHROPIC_AUTH_TOKEN": "<your token>",
    "ANTHROPIC_MODEL": "llama3.1:8b",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "llama3.1:8b"
  }
}
  • ANTHROPIC_MODEL sets the model for the session. ANTHROPIC_DEFAULT_HAIKU_MODEL sets the model for background tasks (it replaces the deprecated ANTHROPIC_SMALL_FAST_MODEL). CLAUDE_CODE_SUBAGENT_MODEL sets the subagent model. Claude Code does not validate non-Claude model names, so use the exact name from /v1/models.
  • The machine's runner must expose an Anthropic-compatible /v1/messages endpoint. Ollama documents Anthropic Messages support (streaming, tool calling; no tool_choice, prompt caching or token counting), so use a recent Ollama. For a runner without it, put a translating proxy such as LiteLLM in front of it and add it with --provider.
  • Token counting is optional: without it Claude Code estimates context use. If a runner rejects newer request fields with a 400, try CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1.
  • Anthropic's documentation states it does not support routing Claude Code to non-Claude models through a gateway, so treat this setup as best effort. Run /status in Claude Code to confirm the base URL and credential in use.

OpenAI Codex CLI

Add a provider to ~/.codex/config.toml and select it. Current Codex documentation says responses is the only supported wire_api and the default, so leave the line out. Older guides that set wire_api = "chat" no longer apply: Chat Completions support in Codex was removed in early 2026.

toml
# ~/.codex/config.toml
model = "llama3.1:8b"
model_provider = "tokmine"

[model_providers.tokmine]
name = "Tokmine"
base_url = "https://tunnel.tokmine.ai/t/<slug>/v1"
env_key = "TOKMINE_TOKEN"

That means the machine's runner must support the OpenAI Responses API (/v1/responses, which Tokmine forwards). Ollama supports it from v0.13.3, without stateful features (previous_response_id, conversation). Provider ids openai, ollama and lmstudio are reserved, so pick another id such as tokmine. Export TOKMINE_TOKEN first; env_key names the variable Codex reads.

OpenCode

Add a custom OpenAI-compatible provider to opencode.json (project root, or your global OpenCode config). {env:VAR} reads an environment variable.

json
{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "tokmine": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "Tokmine",
      "options": {
        "baseURL": "https://tunnel.tokmine.ai/t/<slug>/v1",
        "apiKey": "{env:TOKMINE_TOKEN}"
      },
      "models": {
        "llama3.1:8b": {
          "name": "Llama 3.1 8B",
          "limit": { "context": 32768, "output": 4096 }
        }
      }
    }
  }
}

Aider

Set the two OpenAI variables and prefix the model with openai/. Aider also reads them from a .env file, which you should not commit.

bash
export OPENAI_API_BASE=https://tunnel.tokmine.ai/t/<slug>/v1
export OPENAI_API_KEY=$TOKMINE_TOKEN
aider --model openai/llama3.1:8b

On Windows, setx OPENAI_API_BASE ... and setx OPENAI_API_KEY ..., then restart the shell. Aider may warn about unfamiliar models; that is expected.

Continue

Edit ~/.continue/config.yaml. Use the openai provider with apiBase:

yaml
name: Tokmine
version: 0.0.1
schema: v1

models:
  - name: Tokmine llama3.1
    provider: openai
    model: llama3.1:8b
    apiBase: https://tunnel.tokmine.ai/t/<slug>/v1
    apiKey: <your token>

Continue's docs show the key inline in this file. That file lives in your home folder, not the repository, so keep it there.

Cline and Roo Code

In Cline, choose the OpenAI Compatible provider and fill in three fields: Base URL https://tunnel.tokmine.ai/t/<slug>/v1, your token as API key, and the model ID. Advanced fields (context window, max output tokens, image support) should match your local model. Roo Code is a fork with a similar OpenAI Compatible provider, but I could not verify its current fields: check the tool's docs.

Cursor

Cursor has an Override OpenAI Base URL setting (Settings, Models). I could not confirm its behaviour from Cursor's own documentation, so check the tool's docs. Community reports say Cursor sends these requests from its own servers, not from your machine. That is why localhost addresses fail and a public URL such as a Tokmine endpoint is needed. Reports also say the override applies to chat-style requests with a model name you add yourself, not to every Cursor feature (Tab completion, for example, uses Cursor's own models). Expect differences in agent mode and confirm in the app.

text
Override OpenAI Base URL:  https://tunnel.tokmine.ai/t/<slug>/v1
OpenAI API Key:            <your token>
Custom model name:         llama3.1:8b

Zed

Add an OpenAI-compatible provider in Zed's settings.json:

json
{
  "language_models": {
    "openai_compatible": {
      "tokmine": {
        "api_url": "https://tunnel.tokmine.ai/t/<slug>/v1",
        "available_models": [
          { "name": "llama3.1:8b", "display_name": "Llama 3.1 8B", "max_tokens": 32768 }
        ]
      }
    }
  }
}

Zed reads the key from an environment variable named after the provider id in upper case with _API_KEY appended (here TOKMINE_API_KEY), or you can enter it in the provider settings. Zed warns not to put API keys in settings.json. Tool support is on by default; disable it for models that cannot call tools.

VS Code (Copilot chat, bring your own model)

VS Code can add a Custom Endpoint under Chat: Manage Language Models, then Add Models. Pick the API type (Chat Completions, Responses or Messages) and edit the generated chatLanguageModels.json. The model url is the full endpoint path.

json
[
  {
    "name": "Tokmine",
    "vendor": "customendpoint",
    "apiKey": "${input:tokmineToken}",
    "apiType": "chat-completions",
    "models": [
      {
        "id": "llama3.1:8b",
        "name": "Llama 3.1 8B",
        "url": "https://tunnel.tokmine.ai/t/<slug>/v1/chat/completions",
        "toolCalling": true,
        "maxInputTokens": 32768,
        "maxOutputTokens": 4096
      }
    ]
  }
]

Set toolCalling to true only for models that support tools. Some Copilot features still need a GitHub sign-in.

Open WebUI and LibreChat

Open WebUI: Admin Settings, Connections, add a connection with URL https://tunnel.tokmine.ai/t/<slug>/v1 and your token. Models are detected from /models, or add IDs by hand. Or set environment variables when starting it:

bash
OPENAI_API_BASE_URL=https://tunnel.tokmine.ai/t/<slug>/v1
OPENAI_API_KEY=<your token>

LibreChat: add a custom endpoint to librechat.yaml and set TOKMINE_TOKEN in its .env:

yaml
endpoints:
  custom:
    - name: "Tokmine"
      apiKey: "${TOKMINE_TOKEN}"
      baseURL: "https://tunnel.tokmine.ai/t/<slug>/v1"
      models:
        default: ["llama3.1:8b"]
        fetch: true

LangChain, LlamaIndex, Vercel AI SDK and LiteLLM

LangChain (Python)

python
from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    model="llama3.1:8b",
    base_url="https://tunnel.tokmine.ai/t/<slug>/v1",
    api_key="<your token>",
)
print(llm.invoke("Hello").content)

LangChain warns that ChatOpenAI targets the official OpenAI specification and may drop non-standard fields from other providers.

LlamaIndex

python
# pip install llama-index-llms-openai-like
from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    model="llama3.1:8b",
    api_base="https://tunnel.tokmine.ai/t/<slug>/v1",
    api_key="<your token>",
    context_window=32768,
    is_chat_model=True,
)
print(llm.complete("Hello"))

Vercel AI SDK

javascript
// npm i @ai-sdk/openai-compatible ai
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { generateText } from "ai";

const tokmine = createOpenAICompatible({
  name: "tokmine",
  baseURL: "https://tunnel.tokmine.ai/t/<slug>/v1",
  apiKey: process.env.TOKMINE_TOKEN,
});
const { text } = await generateText({ model: tokmine("llama3.1:8b"), prompt: "Hello" });

LiteLLM

Prefix the model with openai/ and pass the base URL including /v1. LiteLLM docs note that a Not Found error usually means the /v1 is missing.

python
import litellm

resp = litellm.completion(
    model="openai/llama3.1:8b",
    api_base="https://tunnel.tokmine.ai/t/<slug>/v1",
    api_key="<your token>",
    messages=[{"role": "user", "content": "Hello"}],
)
yaml
model_list:
  - model_name: local-llama
    litellm_params:
      model: openai/llama3.1:8b
      api_base: https://tunnel.tokmine.ai/t/<slug>/v1
      api_key: os.environ/TOKMINE_TOKEN

Any OpenAI-compatible client

text
Base URL : https://tunnel.tokmine.ai/t/<slug>/v1
API key  : <your token>       (sent as Authorization: Bearer)
Model    : a name from GET <base>/models, e.g. llama3.1:8b

Streaming, tool calls and embeddings work when the machine's runner supports them. Header names and paths are listed in API compatibility.

Gemini-style clients

Use the base URL without /v1 and send the token as x-goog-api-key (or the SDK's API key setting). The endpoints are /v1beta/models/<model>:generateContent and :streamGenerateContent.

bash
curl "https://tunnel.tokmine.ai/t/<slug>/v1beta/models/llama3.1:8b:generateContent" \
  -H "x-goog-api-key: $TOKMINE_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"contents":[{"parts":[{"text":"Hello"}]}]}'

Caveats

  • Tool calling and agents depend on the model. Pick a local model that supports tool calls; small or non-tool models loop, ignore tools or produce malformed calls in coding agents. Test with a small task first.
  • Context window. Coding tools send a lot of context. Many runners default to a small window (Ollama, for example); raise it on the runner and tell the client the real size where it has a setting.
  • Cold starts. The first request after the model was idle can take a while to load. Streaming requests receive keep-alive comments so clients and CDNs do not time out; non-streaming requests do not.
  • Rate limits. Anonymous sessions are rate limited and one is allowed per IP. Signed-in workspaces get higher limits.
  • Anonymous URLs expire after 30 minutes. For editors and CI create a static endpoint, which is a stable URL that several machines can share.
  • Formats change. These settings were checked against each tool's documentation at the time of writing; tools update often, so check the tool's docs if something is rejected.

See also Troubleshooting for 401, 404 and 502 responses.