Running a model locally is easy. Running it for other people, from other places, on more than one machine, is where it gets awkward.

The single-machine honeymoon

You install Ollama, pull a model, and it answers on localhost:11434. Everything works because the only client is on the same machine.

Then reality shows up

A teammate wants to try the model. Your laptop is behind a router. The good GPU is at home and you are not. A second box appears in the cloud. Now you need to answer questions that a hosted API answers for you:

  • How do people reach it without port forwarding or a VPN?
  • How do we know who is calling it?
  • Which machine should take this request, when one is already busy?
  • What happens when a machine goes to sleep?

A tunnel is not enough

General purpose tunnels solve reachability. They do not know about models, API dialects, or load. If two machines serve the same model, a tunnel has no opinion about which should answer. A traffic manager does: it tracks CPU, RAM, GPU and queue depth for each machine, prefers one that already has the model loaded, skips unhealthy ones, and retries elsewhere if a request fails before it starts.

Keep the SDKs you have

The other half is compatibility. If the public URL speaks OpenAI, Anthropic and Gemini, existing clients, agents and CI jobs work by changing a base URL and a key.

Try it

Run tokmine on a machine with a model server and you get a public URL in seconds. See the quickstart.