[{"data":1,"prerenderedAt":69},["ShallowReactive",2],{"blog-post-ai-coding-tools-on-your-own-gpu":3,"blog-index":16},{"slug":4,"title":5,"excerpt":6,"author":7,"authorBio":8,"publishedAt":9,"readMinutes":10,"tags":11,"coverImage":14,"body":15},"ai-coding-tools-on-your-own-gpu","Point Claude Code, Cline and Aider at your own GPU","Any coding tool that lets you change the base URL can use the models on your own hardware. Here are the three values every tool needs.","The Tokmine team","We build Tokmine, a traffic manager for local AI models.","2026-10-01T09:00:00.000Z",5,[12,13],"Coding tools","Tutorial","\u002Fblog\u002Fai-coding-tools-on-your-own-gpu.svg","\n      \u003Cp class=\"lead\">Editors and terminal agents are moving toward \"bring your own model\". If your model runs on a machine at home, the only missing piece is a URL the tool can reach.\u003C\u002Fp>\n\n      \u003Ch2>Three values, every tool\u003C\u002Fh2>\n      \u003Cp>Whatever the tool, you need the same three facts:\u003C\u002Fp>\n      \u003Cul>\n        \u003Cli>\u003Cstrong>Base URL.\u003C\u002Fstrong> \u003Ccode>https:\u002F\u002Ftunnel.tokmine.ai\u002Ft\u002F&lt;slug&gt;\u003C\u002Fcode>, with \u003Ccode>\u002Fv1\u003C\u002Fcode> added for OpenAI-style clients.\u003C\u002Fli>\n        \u003Cli>\u003Cstrong>Token.\u003C\u002Fstrong> Your own account token from the portal.\u003C\u002Fli>\n        \u003Cli>\u003Cstrong>Model name.\u003C\u002Fstrong> Exactly as the model list returns it, for example \u003Ccode>llama3.1:8b\u003C\u002Fcode>.\u003C\u002Fli>\n      \u003C\u002Ful>\n\n      \u003Ch2>Verify before you configure\u003C\u002Fh2>\n      \u003Cp>Call the model list with your token first. If that works, any later failure is a client setting, not the network.\u003C\u002Fp>\n      \u003Cpre>\u003Ccode>curl https:\u002F\u002Ftunnel.tokmine.ai\u002Ft\u002F&lt;slug&gt;\u002Fv1\u002Fmodels \\\n  -H \"Authorization: Bearer $TOKMINE_TOKEN\"\u003C\u002Fcode>\u003C\u002Fpre>\n\n      \u003Ch2>Use a static endpoint for editor configs\u003C\u002Fh2>\n      \u003Cp>Anonymous session URLs expire after 30 minutes, which is no good in a config file. A static endpoint keeps the same URL, so you set it once.\u003C\u002Fp>\n\n      \u003Ch2>Tool notes\u003C\u002Fh2>\n      \u003Cp>Claude Code speaks the Anthropic Messages API, so use the URL without \u003Ccode>\u002Fv1\u003C\u002Fcode> and a recent Ollama. Anthropic does not support routing Claude Code to non-Claude models, so treat that setup as best effort. Codex needs a runner that supports the OpenAI Responses API. Aider, Continue, Cline and Zed take the OpenAI-style URL with \u003Ccode>\u002Fv1\u003C\u002Fcode>.\u003C\u002Fp>\n\n      \u003Ch2>Full guide\u003C\u002Fh2>\n      \u003Cp>Copy-paste configs for each tool are in \u003Ca href=\"\u002Fdocs\u002Fai-clients\">AI tools and clients\u003C\u002Fa>.\u003C\u002Fp>\n    ",[17,27,35,37,46,54,62],{"slug":18,"title":19,"excerpt":20,"author":7,"authorBio":8,"publishedAt":21,"readMinutes":22,"tags":23,"coverImage":26},"why-tunnels-fall-short-for-local-llms","Why a plain tunnel falls short for local LLMs","A general purpose tunnel gets a request to your machine. It does not know about models, users or load, and those are the parts you end up needing.","2026-10-03T09:00:00.000Z",4,[24,25],"Architecture","Local models","\u002Fblog\u002Fwhy-tunnels-fall-short-for-local-llms.svg",{"slug":28,"title":29,"excerpt":30,"author":7,"authorBio":8,"publishedAt":31,"readMinutes":22,"tags":32,"coverImage":34},"load-balancing-local-llm-machines","Two GPU machines, one endpoint: load balancing local models","Add a second machine and requests go to the one with more headroom. No reverse proxy to build and no client changes when machines come and go.","2026-10-02T09:00:00.000Z",[33,24],"Load balancing","\u002Fblog\u002Fload-balancing-local-llm-machines.svg",{"slug":4,"title":5,"excerpt":6,"author":7,"authorBio":8,"publishedAt":9,"readMinutes":10,"tags":36,"coverImage":14},[12,13],{"slug":38,"title":39,"excerpt":40,"author":7,"authorBio":8,"publishedAt":41,"readMinutes":22,"tags":42,"coverImage":45},"one-token-per-teammate","One token per teammate: share a local model safely","Passing around one API key works until it leaks or someone runs a job all night. Per-user tokens make a shared model safe to hand out.","2026-09-30T09:00:00.000Z",[43,44],"Teams","Security","\u002Fblog\u002Fone-token-per-teammate.svg",{"slug":47,"title":48,"excerpt":49,"author":7,"authorBio":8,"publishedAt":50,"readMinutes":10,"tags":51,"coverImage":53},"local-vs-hosted-ai-cost","Local or hosted AI? What running your own models costs","Hosted APIs bill per token. Local models bill in hardware, power and your time. Here is how to compare the two without fooling yourself.","2026-09-29T09:00:00.000Z",[52,25],"Cost","\u002Fblog\u002Flocal-vs-hosted-ai-cost.svg",{"slug":55,"title":56,"excerpt":57,"author":7,"authorBio":8,"publishedAt":58,"readMinutes":22,"tags":59,"coverImage":61},"use-ollama-from-anywhere","Use Ollama from anywhere, in three steps","Reach the models on your home GPU from a laptop, a CI job or a teammate's machine, without opening a single port or setting up a VPN.","2026-09-22T09:00:00.000Z",[60,13],"Ollama","\u002Fblog\u002Fuse-ollama-from-anywhere.svg",{"slug":63,"title":64,"excerpt":65,"author":7,"authorBio":8,"publishedAt":66,"readMinutes":10,"tags":67,"coverImage":68},"why-local-models-need-a-traffic-manager","Why local AI models need a traffic manager","One GPU box is a hobby. Three machines and five teammates is a routing problem. Here is what changes when your models live on your own hardware.","2026-09-15T09:00:00.000Z",[25,24],"\u002Fblog\u002Fwhy-local-models-need-a-traffic-manager.svg",1791297637589]