Use your Tokmine endpoint in AI coding tools and clients
Any tool that lets you change the API base URL can talk to your local models through Tokmine. Every client needs the same three facts: a base URL, a token and a model name.
Find your values
| Base URL | https://tunnel.tokmine.ai/t/<slug>. Clients that speak the OpenAI API want /v1 added: https://tunnel.tokmine.ai/t/<slug>/v1. Anthropic and Gemini clients use the URL without /v1. |
|---|---|
| Token | Anonymous session: the token the CLI prints. Signed in: your own account token from the portal. Not the agent token. |
| Model name | Exactly as returned by the model list, for example llama3.1:8b. |
The URL is printed by tokmine when it connects, and shown with copy buttons on the Endpoints page of the portal. For editor and tool configs use a static endpoint (https://tunnel.tokmine.ai/t/<slug> that never changes), because anonymous session URLs expire after 30 minutes.
Verify before configuring a tool
export TOKMINE_TOKEN=<your token>
curl https://tunnel.tokmine.ai/t/<slug>/v1/models \
-H "Authorization: Bearer $TOKMINE_TOKEN" The response lists the models your online machines serve. If that works, any failure in a tool is a client setting. Use these values in the sections below (replace <slug>).
TOKMINE_TOKEN wherever the tool allows it. On this page
- Claude Code, OpenAI Codex CLI, OpenCode, Aider
- Continue, Cline and Roo Code, Cursor, Zed, VS Code
- Open WebUI and LibreChat, LangChain, LlamaIndex, Vercel AI SDK, LiteLLM
- Any OpenAI-compatible client, Gemini-style clients, Caveats
Claude Code
Claude Code speaks the Anthropic Messages API. Point it at your URL without/v1; it appends /v1/messages itself. Anthropic's gateway documentation lists ANTHROPIC_AUTH_TOKEN (sent as Authorization: Bearer) and ANTHROPIC_API_KEY (sent as x-api-key); Tokmine accepts either header.
export ANTHROPIC_BASE_URL=https://tunnel.tokmine.ai/t/<slug>
export ANTHROPIC_AUTH_TOKEN=$TOKMINE_TOKEN
export ANTHROPIC_MODEL=llama3.1:8b
export ANTHROPIC_DEFAULT_HAIKU_MODEL=llama3.1:8b
claudeOr persist it in the env block of ~/.claude/settings.json (%USERPROFILE%\.claude\settings.json on Windows):
{
"env": {
"ANTHROPIC_BASE_URL": "https://tunnel.tokmine.ai/t/<slug>",
"ANTHROPIC_AUTH_TOKEN": "<your token>",
"ANTHROPIC_MODEL": "llama3.1:8b",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "llama3.1:8b"
}
}ANTHROPIC_MODELsets the model for the session.ANTHROPIC_DEFAULT_HAIKU_MODELsets the model for background tasks (it replaces the deprecatedANTHROPIC_SMALL_FAST_MODEL).CLAUDE_CODE_SUBAGENT_MODELsets the subagent model. Claude Code does not validate non-Claude model names, so use the exact name from/v1/models.- The machine's runner must expose an Anthropic-compatible
/v1/messagesendpoint. Ollama documents Anthropic Messages support (streaming, tool calling; notool_choice, prompt caching or token counting), so use a recent Ollama. For a runner without it, put a translating proxy such as LiteLLM in front of it and add it with--provider. - Token counting is optional: without it Claude Code estimates context use. If a runner rejects newer request fields with a
400, tryCLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS=1. - Anthropic's documentation states it does not support routing Claude Code to non-Claude models through a gateway, so treat this setup as best effort. Run
/statusin Claude Code to confirm the base URL and credential in use.
OpenAI Codex CLI
Add a provider to ~/.codex/config.toml and select it. Current Codex documentation says responses is the only supported wire_api and the default, so leave the line out. Older guides that set wire_api = "chat" no longer apply: Chat Completions support in Codex was removed in early 2026.
# ~/.codex/config.toml
model = "llama3.1:8b"
model_provider = "tokmine"
[model_providers.tokmine]
name = "Tokmine"
base_url = "https://tunnel.tokmine.ai/t/<slug>/v1"
env_key = "TOKMINE_TOKEN" That means the machine's runner must support the OpenAI Responses API (/v1/responses, which Tokmine forwards). Ollama supports it from v0.13.3, without stateful features (previous_response_id, conversation). Provider ids openai, ollama and lmstudio are reserved, so pick another id such as tokmine. Export TOKMINE_TOKEN first; env_key names the variable Codex reads.
OpenCode
Add a custom OpenAI-compatible provider to opencode.json (project root, or your global OpenCode config). {env:VAR} reads an environment variable.
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"tokmine": {
"npm": "@ai-sdk/openai-compatible",
"name": "Tokmine",
"options": {
"baseURL": "https://tunnel.tokmine.ai/t/<slug>/v1",
"apiKey": "{env:TOKMINE_TOKEN}"
},
"models": {
"llama3.1:8b": {
"name": "Llama 3.1 8B",
"limit": { "context": 32768, "output": 4096 }
}
}
}
}
}Aider
Set the two OpenAI variables and prefix the model with openai/. Aider also reads them from a .env file, which you should not commit.
export OPENAI_API_BASE=https://tunnel.tokmine.ai/t/<slug>/v1
export OPENAI_API_KEY=$TOKMINE_TOKEN
aider --model openai/llama3.1:8bOn Windows, setx OPENAI_API_BASE ... and setx OPENAI_API_KEY ..., then restart the shell. Aider may warn about unfamiliar models; that is expected.
Continue
Edit ~/.continue/config.yaml. Use the openai provider with apiBase:
name: Tokmine
version: 0.0.1
schema: v1
models:
- name: Tokmine llama3.1
provider: openai
model: llama3.1:8b
apiBase: https://tunnel.tokmine.ai/t/<slug>/v1
apiKey: <your token>Continue's docs show the key inline in this file. That file lives in your home folder, not the repository, so keep it there.
Cline and Roo Code
In Cline, choose the OpenAI Compatible provider and fill in three fields: Base URL https://tunnel.tokmine.ai/t/<slug>/v1, your token as API key, and the model ID. Advanced fields (context window, max output tokens, image support) should match your local model. Roo Code is a fork with a similar OpenAI Compatible provider, but I could not verify its current fields: check the tool's docs.
Cursor
Cursor has an Override OpenAI Base URL setting (Settings, Models). I could not confirm its behaviour from Cursor's own documentation, so check the tool's docs. Community reports say Cursor sends these requests from its own servers, not from your machine. That is why localhost addresses fail and a public URL such as a Tokmine endpoint is needed. Reports also say the override applies to chat-style requests with a model name you add yourself, not to every Cursor feature (Tab completion, for example, uses Cursor's own models). Expect differences in agent mode and confirm in the app.
Override OpenAI Base URL: https://tunnel.tokmine.ai/t/<slug>/v1
OpenAI API Key: <your token>
Custom model name: llama3.1:8bZed
Add an OpenAI-compatible provider in Zed's settings.json:
{
"language_models": {
"openai_compatible": {
"tokmine": {
"api_url": "https://tunnel.tokmine.ai/t/<slug>/v1",
"available_models": [
{ "name": "llama3.1:8b", "display_name": "Llama 3.1 8B", "max_tokens": 32768 }
]
}
}
}
} Zed reads the key from an environment variable named after the provider id in upper case with _API_KEY appended (here TOKMINE_API_KEY), or you can enter it in the provider settings. Zed warns not to put API keys in settings.json. Tool support is on by default; disable it for models that cannot call tools.
VS Code (Copilot chat, bring your own model)
VS Code can add a Custom Endpoint under Chat: Manage Language Models, then Add Models. Pick the API type (Chat Completions, Responses or Messages) and edit the generated chatLanguageModels.json. The model url is the full endpoint path.
[
{
"name": "Tokmine",
"vendor": "customendpoint",
"apiKey": "${input:tokmineToken}",
"apiType": "chat-completions",
"models": [
{
"id": "llama3.1:8b",
"name": "Llama 3.1 8B",
"url": "https://tunnel.tokmine.ai/t/<slug>/v1/chat/completions",
"toolCalling": true,
"maxInputTokens": 32768,
"maxOutputTokens": 4096
}
]
}
]Set toolCalling to true only for models that support tools. Some Copilot features still need a GitHub sign-in.
Open WebUI and LibreChat
Open WebUI: Admin Settings, Connections, add a connection with URL https://tunnel.tokmine.ai/t/<slug>/v1 and your token. Models are detected from /models, or add IDs by hand. Or set environment variables when starting it:
OPENAI_API_BASE_URL=https://tunnel.tokmine.ai/t/<slug>/v1
OPENAI_API_KEY=<your token>LibreChat: add a custom endpoint to librechat.yaml and set TOKMINE_TOKEN in its .env:
endpoints:
custom:
- name: "Tokmine"
apiKey: "${TOKMINE_TOKEN}"
baseURL: "https://tunnel.tokmine.ai/t/<slug>/v1"
models:
default: ["llama3.1:8b"]
fetch: trueLangChain, LlamaIndex, Vercel AI SDK and LiteLLM
LangChain (Python)
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
model="llama3.1:8b",
base_url="https://tunnel.tokmine.ai/t/<slug>/v1",
api_key="<your token>",
)
print(llm.invoke("Hello").content)LangChain warns that ChatOpenAI targets the official OpenAI specification and may drop non-standard fields from other providers.
LlamaIndex
# pip install llama-index-llms-openai-like
from llama_index.llms.openai_like import OpenAILike
llm = OpenAILike(
model="llama3.1:8b",
api_base="https://tunnel.tokmine.ai/t/<slug>/v1",
api_key="<your token>",
context_window=32768,
is_chat_model=True,
)
print(llm.complete("Hello"))Vercel AI SDK
// npm i @ai-sdk/openai-compatible ai
import { createOpenAICompatible } from "@ai-sdk/openai-compatible";
import { generateText } from "ai";
const tokmine = createOpenAICompatible({
name: "tokmine",
baseURL: "https://tunnel.tokmine.ai/t/<slug>/v1",
apiKey: process.env.TOKMINE_TOKEN,
});
const { text } = await generateText({ model: tokmine("llama3.1:8b"), prompt: "Hello" });LiteLLM
Prefix the model with openai/ and pass the base URL including /v1. LiteLLM docs note that a Not Found error usually means the /v1 is missing.
import litellm
resp = litellm.completion(
model="openai/llama3.1:8b",
api_base="https://tunnel.tokmine.ai/t/<slug>/v1",
api_key="<your token>",
messages=[{"role": "user", "content": "Hello"}],
)model_list:
- model_name: local-llama
litellm_params:
model: openai/llama3.1:8b
api_base: https://tunnel.tokmine.ai/t/<slug>/v1
api_key: os.environ/TOKMINE_TOKENAny OpenAI-compatible client
Base URL : https://tunnel.tokmine.ai/t/<slug>/v1
API key : <your token> (sent as Authorization: Bearer)
Model : a name from GET <base>/models, e.g. llama3.1:8bStreaming, tool calls and embeddings work when the machine's runner supports them. Header names and paths are listed in API compatibility.
Gemini-style clients
Use the base URL without /v1 and send the token as x-goog-api-key (or the SDK's API key setting). The endpoints are /v1beta/models/<model>:generateContent and :streamGenerateContent.
curl "https://tunnel.tokmine.ai/t/<slug>/v1beta/models/llama3.1:8b:generateContent" \
-H "x-goog-api-key: $TOKMINE_TOKEN" \
-H "Content-Type: application/json" \
-d '{"contents":[{"parts":[{"text":"Hello"}]}]}'Caveats
- Tool calling and agents depend on the model. Pick a local model that supports tool calls; small or non-tool models loop, ignore tools or produce malformed calls in coding agents. Test with a small task first.
- Context window. Coding tools send a lot of context. Many runners default to a small window (Ollama, for example); raise it on the runner and tell the client the real size where it has a setting.
- Cold starts. The first request after the model was idle can take a while to load. Streaming requests receive keep-alive comments so clients and CDNs do not time out; non-streaming requests do not.
- Rate limits. Anonymous sessions are rate limited and one is allowed per IP. Signed-in workspaces get higher limits.
- Anonymous URLs expire after 30 minutes. For editors and CI create a static endpoint, which is a stable URL that several machines can share.
- Formats change. These settings were checked against each tool's documentation at the time of writing; tools update often, so check the tool's docs if something is rejected.
See also Troubleshooting for 401, 404 and 502 responses.