[{"data":1,"prerenderedAt":69},["ShallowReactive",2],{"blog-post-load-balancing-local-llm-machines":3,"blog-index":16},{"slug":4,"title":5,"excerpt":6,"author":7,"authorBio":8,"publishedAt":9,"readMinutes":10,"tags":11,"coverImage":14,"body":15},"load-balancing-local-llm-machines","Two GPU machines, one endpoint: load balancing local models","Add a second machine and requests go to the one with more headroom. No reverse proxy to build and no client changes when machines come and go.","The Tokmine team","We build Tokmine, a traffic manager for local AI models.","2026-10-02T09:00:00.000Z",4,[12,13],"Load balancing","Architecture","\u002Fblog\u002Fload-balancing-local-llm-machines.svg","\n      \u003Cp class=\"lead\">One GPU box eventually gets busy. The usual fix, a second box, creates a new question: which one answers?\u003C\u002Fp>\n\n      \u003Ch2>The DIY version\u003C\u002Fh2>\n      \u003Cp>You can put a reverse proxy in front of two machines and write rules for health checks, retries and stickiness. It works, and it becomes a small project you now maintain, plus you still need a way to reach both from outside your network.\u003C\u002Fp>\n\n      \u003Ch2>One static endpoint, several machines\u003C\u002Fh2>\n      \u003Cp>With a static endpoint, a named URL such as \u003Ccode>https:\u002F\u002Ftunnel.tokmine.ai\u002Ft\u002Facme-prod\u003C\u002Fcode> stays the same and several of your machines can share it. Start each machine against it:\u003C\u002Fp>\n      \u003Cpre>\u003Ccode>tokmine --endpoint acme-prod\u003C\u002Fcode>\u003C\u002Fpre>\n\n      \u003Ch2>How requests are routed\u003C\u002Fh2>\n      \u003Cul>\n        \u003Cli>Requests are routed by model name. The model must be served by at least one online machine.\u003C\u002Fli>\n        \u003Cli>When several machines serve it, the request goes to the one with the most headroom.\u003C\u002Fli>\n        \u003Cli>A machine that goes offline simply stops receiving traffic.\u003C\u002Fli>\n      \u003C\u002Ful>\n      \u003Cp>Clients see one URL and one token. Adding or removing a machine needs no client change.\u003C\u002Fp>\n\n      \u003Ch2>Targeting a specific machine\u003C\u002Fh2>\n      \u003Cp>Sometimes you want a particular provider, for example to test a new model on one box. Add \u003Ccode>\u002Fp\u002F&lt;providerId&gt;\u003C\u002Fcode> after the slug to pin the request.\u003C\u002Fp>\n\n      \u003Ch2>Try it\u003C\u002Fh2>\n      \u003Cp>Static endpoints are for paid accounts. Setup steps are in the \u003Ca href=\"\u002Fdocs\u002Fendpoints\">endpoints docs\u003C\u002Fa>.\u003C\u002Fp>\n    ",[17,25,27,37,46,54,62],{"slug":18,"title":19,"excerpt":20,"author":7,"authorBio":8,"publishedAt":21,"readMinutes":10,"tags":22,"coverImage":24},"why-tunnels-fall-short-for-local-llms","Why a plain tunnel falls short for local LLMs","A general purpose tunnel gets a request to your machine. It does not know about models, users or load, and those are the parts you end up needing.","2026-10-03T09:00:00.000Z",[13,23],"Local models","\u002Fblog\u002Fwhy-tunnels-fall-short-for-local-llms.svg",{"slug":4,"title":5,"excerpt":6,"author":7,"authorBio":8,"publishedAt":9,"readMinutes":10,"tags":26,"coverImage":14},[12,13],{"slug":28,"title":29,"excerpt":30,"author":7,"authorBio":8,"publishedAt":31,"readMinutes":32,"tags":33,"coverImage":36},"ai-coding-tools-on-your-own-gpu","Point Claude Code, Cline and Aider at your own GPU","Any coding tool that lets you change the base URL can use the models on your own hardware. Here are the three values every tool needs.","2026-10-01T09:00:00.000Z",5,[34,35],"Coding tools","Tutorial","\u002Fblog\u002Fai-coding-tools-on-your-own-gpu.svg",{"slug":38,"title":39,"excerpt":40,"author":7,"authorBio":8,"publishedAt":41,"readMinutes":10,"tags":42,"coverImage":45},"one-token-per-teammate","One token per teammate: share a local model safely","Passing around one API key works until it leaks or someone runs a job all night. Per-user tokens make a shared model safe to hand out.","2026-09-30T09:00:00.000Z",[43,44],"Teams","Security","\u002Fblog\u002Fone-token-per-teammate.svg",{"slug":47,"title":48,"excerpt":49,"author":7,"authorBio":8,"publishedAt":50,"readMinutes":32,"tags":51,"coverImage":53},"local-vs-hosted-ai-cost","Local or hosted AI? What running your own models costs","Hosted APIs bill per token. Local models bill in hardware, power and your time. Here is how to compare the two without fooling yourself.","2026-09-29T09:00:00.000Z",[52,23],"Cost","\u002Fblog\u002Flocal-vs-hosted-ai-cost.svg",{"slug":55,"title":56,"excerpt":57,"author":7,"authorBio":8,"publishedAt":58,"readMinutes":10,"tags":59,"coverImage":61},"use-ollama-from-anywhere","Use Ollama from anywhere, in three steps","Reach the models on your home GPU from a laptop, a CI job or a teammate's machine, without opening a single port or setting up a VPN.","2026-09-22T09:00:00.000Z",[60,35],"Ollama","\u002Fblog\u002Fuse-ollama-from-anywhere.svg",{"slug":63,"title":64,"excerpt":65,"author":7,"authorBio":8,"publishedAt":66,"readMinutes":32,"tags":67,"coverImage":68},"why-local-models-need-a-traffic-manager","Why local AI models need a traffic manager","One GPU box is a hobby. Three machines and five teammates is a routing problem. Here is what changes when your models live on your own hardware.","2026-09-15T09:00:00.000Z",[23,13],"\u002Fblog\u002Fwhy-local-models-need-a-traffic-manager.svg",1791297637594]