"Local is free" and "hosted is cheaper than a GPU" are both true, for different people. The honest answer depends on how much you use the model and how many people share it.
What hosted APIs charge for
A hosted API bills per token, so cost scales with usage. Zero usage costs nothing, and a busy coding agent or batch job can turn into a surprise invoice. You pay for the convenience of never owning a machine.
What local models charge for
- Hardware. A GPU you may already own for gaming or work.
- Power. Small when idle, real under sustained load.
- Your time. Installing runners, pulling models, keeping things updated.
- Access. Making the model reachable from where you actually work.
The first item is often already paid for. The last one is the hidden cost, and it is the one that makes people give up on local models after a week.
A simple way to compare
Take a normal week and count the tokens your tools send. Multiply by the hosted price for the model class you would otherwise use. Compare that with your hardware cost spread over its useful life, plus electricity. If a model on hardware you already own is good enough for the task, local wins on cost. If you need a frontier model, it does not, and that is fine.
Where local models fit best
High-volume, lower-stakes work: autocomplete, summarising, embeddings, test data, drafts, and anything with private code or documents. Keep the hosted model for the hard problems. Nothing forces you to pick one.
The access problem
A model on a home GPU is worth little when you are on a laptop somewhere else. Tokmine gives that machine a public URL with per-user tokens, so the model you already paid for is usable from anywhere. Pricing is $0.50 per machine per month plus $0.50 per person who calls it, listed on the pricing page.
Try it
Run tokmine next to your model server and check the numbers against your own week. The quickstart takes a few minutes.