[{"data":1,"prerenderedAt":46},["ShallowReactive",2],{"blog-post-local-vs-hosted-ai-cost":3,"blog-index":16},{"slug":4,"title":5,"excerpt":6,"author":7,"authorBio":8,"publishedAt":9,"readMinutes":10,"tags":11,"coverImage":14,"body":15},"local-vs-hosted-ai-cost","Local or hosted AI? What running your own models costs","Hosted APIs bill per token. Local models bill in hardware, power and your time. Here is how to compare the two without fooling yourself.","The Tokmine team","We build Tokmine, a traffic manager for local AI models.","2026-09-29T09:00:00.000Z",5,[12,13],"Cost","Local models","\u002Fblog\u002Flocal-vs-hosted-ai-cost.svg","\n      \u003Cp class=\"lead\">\"Local is free\" and \"hosted is cheaper than a GPU\" are both true, for different people. The honest answer depends on how much you use the model and how many people share it.\u003C\u002Fp>\n\n      \u003Ch2>What hosted APIs charge for\u003C\u002Fh2>\n      \u003Cp>A hosted API bills per token, so cost scales with usage. Zero usage costs nothing, and a busy coding agent or batch job can turn into a surprise invoice. You pay for the convenience of never owning a machine.\u003C\u002Fp>\n\n      \u003Ch2>What local models charge for\u003C\u002Fh2>\n      \u003Cul>\n        \u003Cli>\u003Cstrong>Hardware.\u003C\u002Fstrong> A GPU you may already own for gaming or work.\u003C\u002Fli>\n        \u003Cli>\u003Cstrong>Power.\u003C\u002Fstrong> Small when idle, real under sustained load.\u003C\u002Fli>\n        \u003Cli>\u003Cstrong>Your time.\u003C\u002Fstrong> Installing runners, pulling models, keeping things updated.\u003C\u002Fli>\n        \u003Cli>\u003Cstrong>Access.\u003C\u002Fstrong> Making the model reachable from where you actually work.\u003C\u002Fli>\n      \u003C\u002Ful>\n      \u003Cp>The first item is often already paid for. The last one is the hidden cost, and it is the one that makes people give up on local models after a week.\u003C\u002Fp>\n\n      \u003Ch2>A simple way to compare\u003C\u002Fh2>\n      \u003Cp>Take a normal week and count the tokens your tools send. Multiply by the hosted price for the model class you would otherwise use. Compare that with your hardware cost spread over its useful life, plus electricity. If a model on hardware you already own is good enough for the task, local wins on cost. If you need a frontier model, it does not, and that is fine.\u003C\u002Fp>\n\n      \u003Ch2>Where local models fit best\u003C\u002Fh2>\n      \u003Cp>High-volume, lower-stakes work: autocomplete, summarising, embeddings, test data, drafts, and anything with private code or documents. Keep the hosted model for the hard problems. Nothing forces you to pick one.\u003C\u002Fp>\n\n      \u003Ch2>The access problem\u003C\u002Fh2>\n      \u003Cp>A model on a home GPU is worth little when you are on a laptop somewhere else. Tokmine gives that machine a public URL with per-user tokens, so the model you already paid for is usable from anywhere. Pricing is $0.50 per machine per month plus $0.50 per person who calls it, listed on the \u003Ca href=\"\u002Fpricing\">pricing page\u003C\u002Fa>.\u003C\u002Fp>\n\n      \u003Ch2>Try it\u003C\u002Fh2>\n      \u003Cp>Run \u003Ccode>tokmine\u003C\u002Fcode> next to your model server and check the numbers against your own week. The \u003Ca href=\"\u002Fdocs\u002Fquickstart\">quickstart\u003C\u002Fa> takes a few minutes.\u003C\u002Fp>\n    ",[17,27,29,38],{"slug":18,"title":19,"excerpt":20,"author":7,"authorBio":8,"publishedAt":21,"readMinutes":22,"tags":23,"coverImage":26},"one-token-per-teammate","One token per teammate: share a local model safely","Passing around one API key works until it leaks or someone runs a job all night. Per-user tokens make a shared model safe to hand out.","2026-09-30T09:00:00.000Z",4,[24,25],"Teams","Security","\u002Fblog\u002Fone-token-per-teammate.svg",{"slug":4,"title":5,"excerpt":6,"author":7,"authorBio":8,"publishedAt":9,"readMinutes":10,"tags":28,"coverImage":14},[12,13],{"slug":30,"title":31,"excerpt":32,"author":7,"authorBio":8,"publishedAt":33,"readMinutes":22,"tags":34,"coverImage":37},"use-ollama-from-anywhere","Use Ollama from anywhere, in three steps","Reach the models on your home GPU from a laptop, a CI job or a teammate's machine, without opening a single port or setting up a VPN.","2026-09-22T09:00:00.000Z",[35,36],"Ollama","Tutorial","\u002Fblog\u002Fuse-ollama-from-anywhere.svg",{"slug":39,"title":40,"excerpt":41,"author":7,"authorBio":8,"publishedAt":42,"readMinutes":10,"tags":43,"coverImage":45},"why-local-models-need-a-traffic-manager","Why local AI models need a traffic manager","One GPU box is a hobby. Three machines and five teammates is a routing problem. Here is what changes when your models live on your own hardware.","2026-09-15T09:00:00.000Z",[13,44],"Architecture","\u002Fblog\u002Fwhy-local-models-need-a-traffic-manager.svg",1790768272278]