[{"data":1,"prerenderedAt":69},["ShallowReactive",2],{"blog-post-why-tunnels-fall-short-for-local-llms":3,"blog-index":16},{"slug":4,"title":5,"excerpt":6,"author":7,"authorBio":8,"publishedAt":9,"readMinutes":10,"tags":11,"coverImage":14,"body":15},"why-tunnels-fall-short-for-local-llms","Why a plain tunnel falls short for local LLMs","A general purpose tunnel gets a request to your machine. It does not know about models, users or load, and those are the parts you end up needing.","The Tokmine team","We build Tokmine, a traffic manager for local AI models.","2026-10-03T09:00:00.000Z",4,[12,13],"Architecture","Local models","\u002Fblog\u002Fwhy-tunnels-fall-short-for-local-llms.svg","\n      \u003Cp class=\"lead\">Tunnels are great tools. Exposing \u003Ccode>localhost:11434\u003C\u002Fcode> with one command is how many people first use a model away from home. The trouble starts on day two.\u003C\u002Fp>\n\n      \u003Ch2>Where the gaps show up\u003C\u002Fh2>\n      \u003Cul>\n        \u003Cli>\u003Cstrong>Changing URLs.\u003C\u002Fstrong> If the address changes on restart, every config that used it breaks.\u003C\u002Fli>\n        \u003Cli>\u003Cstrong>Session limits.\u003C\u002Fstrong> A tunnel that closes after a while is a poor home for a model your editor calls all day.\u003C\u002Fli>\n        \u003Cli>\u003Cstrong>No auth.\u003C\u002Fstrong> An open URL in front of a GPU is an invitation. Adding auth means adding another layer.\u003C\u002Fli>\n        \u003Cli>\u003Cstrong>No visibility.\u003C\u002Fstrong> You cannot see who called the model or how much.\u003C\u002Fli>\n        \u003Cli>\u003Cstrong>No model awareness.\u003C\u002Fstrong> The tunnel forwards bytes. It does not know which API dialect a client speaks or which machine serves which model.\u003C\u002Fli>\n      \u003C\u002Ful>\n\n      \u003Ch2>What a model-aware layer adds\u003C\u002Fh2>\n      \u003Cp>Tokmine finds the model servers on the machine, detects which API dialects they speak, and puts one URL in front of them that answers OpenAI, Anthropic and Gemini clients. It keeps the same URL when the machine reconnects, requires a token on every request, and shows usage on a dashboard.\u003C\u002Fp>\n\n      \u003Ch2>When a plain tunnel is enough\u003C\u002Fh2>\n      \u003Cp>If it is a quick demo, one person, one afternoon, a simple tunnel is fine. Anonymous Tokmine sessions cover that case too, for 30 minutes without an account.\u003C\u002Fp>\n\n      \u003Ch2>Try it\u003C\u002Fh2>\n      \u003Cp>Run \u003Ccode>tokmine\u003C\u002Fcode> and compare. The \u003Ca href=\"\u002Fdocs\u002Fquickstart\">quickstart\u003C\u002Fa> shows the whole flow.\u003C\u002Fp>\n    ",[17,19,27,37,46,54,62],{"slug":4,"title":5,"excerpt":6,"author":7,"authorBio":8,"publishedAt":9,"readMinutes":10,"tags":18,"coverImage":14},[12,13],{"slug":20,"title":21,"excerpt":22,"author":7,"authorBio":8,"publishedAt":23,"readMinutes":10,"tags":24,"coverImage":26},"load-balancing-local-llm-machines","Two GPU machines, one endpoint: load balancing local models","Add a second machine and requests go to the one with more headroom. No reverse proxy to build and no client changes when machines come and go.","2026-10-02T09:00:00.000Z",[25,12],"Load balancing","\u002Fblog\u002Fload-balancing-local-llm-machines.svg",{"slug":28,"title":29,"excerpt":30,"author":7,"authorBio":8,"publishedAt":31,"readMinutes":32,"tags":33,"coverImage":36},"ai-coding-tools-on-your-own-gpu","Point Claude Code, Cline and Aider at your own GPU","Any coding tool that lets you change the base URL can use the models on your own hardware. Here are the three values every tool needs.","2026-10-01T09:00:00.000Z",5,[34,35],"Coding tools","Tutorial","\u002Fblog\u002Fai-coding-tools-on-your-own-gpu.svg",{"slug":38,"title":39,"excerpt":40,"author":7,"authorBio":8,"publishedAt":41,"readMinutes":10,"tags":42,"coverImage":45},"one-token-per-teammate","One token per teammate: share a local model safely","Passing around one API key works until it leaks or someone runs a job all night. Per-user tokens make a shared model safe to hand out.","2026-09-30T09:00:00.000Z",[43,44],"Teams","Security","\u002Fblog\u002Fone-token-per-teammate.svg",{"slug":47,"title":48,"excerpt":49,"author":7,"authorBio":8,"publishedAt":50,"readMinutes":32,"tags":51,"coverImage":53},"local-vs-hosted-ai-cost","Local or hosted AI? What running your own models costs","Hosted APIs bill per token. Local models bill in hardware, power and your time. Here is how to compare the two without fooling yourself.","2026-09-29T09:00:00.000Z",[52,13],"Cost","\u002Fblog\u002Flocal-vs-hosted-ai-cost.svg",{"slug":55,"title":56,"excerpt":57,"author":7,"authorBio":8,"publishedAt":58,"readMinutes":10,"tags":59,"coverImage":61},"use-ollama-from-anywhere","Use Ollama from anywhere, in three steps","Reach the models on your home GPU from a laptop, a CI job or a teammate's machine, without opening a single port or setting up a VPN.","2026-09-22T09:00:00.000Z",[60,35],"Ollama","\u002Fblog\u002Fuse-ollama-from-anywhere.svg",{"slug":63,"title":64,"excerpt":65,"author":7,"authorBio":8,"publishedAt":66,"readMinutes":32,"tags":67,"coverImage":68},"why-local-models-need-a-traffic-manager","Why local AI models need a traffic manager","One GPU box is a hobby. Three machines and five teammates is a routing problem. Here is what changes when your models live on your own hardware.","2026-09-15T09:00:00.000Z",[13,12],"\u002Fblog\u002Fwhy-local-models-need-a-traffic-manager.svg",1791297637582]