[{"data":1,"prerenderedAt":29},["ShallowReactive",2],{"blog-post-why-local-models-need-a-traffic-manager":3,"blog-index":16},{"slug":4,"title":5,"excerpt":6,"author":7,"authorBio":8,"publishedAt":9,"readMinutes":10,"tags":11,"coverImage":14,"body":15},"why-local-models-need-a-traffic-manager","Why local AI models need a traffic manager","One GPU box is a hobby. Three machines and five teammates is a routing problem. Here is what changes when your models live on your own hardware.","The Tokmine team","We build Tokmine, a traffic manager for local AI models.","2026-09-15T09:00:00.000Z",5,[12,13],"Local models","Architecture","\u002Fblog\u002Fwhy-local-models-need-a-traffic-manager.svg","\n      \u003Cp class=\"lead\">Running a model locally is easy. Running it for other people, from other places, on more than one machine, is where it gets awkward.\u003C\u002Fp>\n\n      \u003Ch2>The single-machine honeymoon\u003C\u002Fh2>\n      \u003Cp>You install Ollama, pull a model, and it answers on \u003Ccode>localhost:11434\u003C\u002Fcode>. Everything works because the only client is on the same machine.\u003C\u002Fp>\n\n      \u003Ch2>Then reality shows up\u003C\u002Fh2>\n      \u003Cp>A teammate wants to try the model. Your laptop is behind a router. The good GPU is at home and you are not. A second box appears in the cloud. Now you need to answer questions that a hosted API answers for you:\u003C\u002Fp>\n      \u003Cul>\n        \u003Cli>How do people reach it without port forwarding or a VPN?\u003C\u002Fli>\n        \u003Cli>How do we know who is calling it?\u003C\u002Fli>\n        \u003Cli>Which machine should take this request, when one is already busy?\u003C\u002Fli>\n        \u003Cli>What happens when a machine goes to sleep?\u003C\u002Fli>\n      \u003C\u002Ful>\n\n      \u003Ch2>A tunnel is not enough\u003C\u002Fh2>\n      \u003Cp>General purpose tunnels solve reachability. They do not know about models, API dialects, or load. If two machines serve the same model, a tunnel has no opinion about which should answer. A traffic manager does: it tracks CPU, RAM, GPU and queue depth for each machine, prefers one that already has the model loaded, skips unhealthy ones, and retries elsewhere if a request fails before it starts.\u003C\u002Fp>\n\n      \u003Ch2>Keep the SDKs you have\u003C\u002Fh2>\n      \u003Cp>The other half is compatibility. If the public URL speaks OpenAI, Anthropic and Gemini, existing clients, agents and CI jobs work by changing a base URL and a key.\u003C\u002Fp>\n\n      \u003Ch2>Try it\u003C\u002Fh2>\n      \u003Cp>Run \u003Ccode>tokmine\u003C\u002Fcode> on a machine with a model server and you get a public URL in seconds. See the \u003Ca href=\"\u002Fdocs\u002Fquickstart\">quickstart\u003C\u002Fa>.\u003C\u002Fp>\n    ",[17,27],{"slug":18,"title":19,"excerpt":20,"author":7,"authorBio":8,"publishedAt":21,"readMinutes":22,"tags":23,"coverImage":26},"use-ollama-from-anywhere","Use Ollama from anywhere, in three steps","Reach the models on your home GPU from a laptop, a CI job or a teammate's machine, without opening a single port or setting up a VPN.","2026-09-22T09:00:00.000Z",4,[24,25],"Ollama","Tutorial","\u002Fblog\u002Fuse-ollama-from-anywhere.svg",{"slug":4,"title":5,"excerpt":6,"author":7,"authorBio":8,"publishedAt":9,"readMinutes":10,"tags":28,"coverImage":14},[12,13],1790694341180]