One GPU box eventually gets busy. The usual fix, a second box, creates a new question: which one answers?

The DIY version

You can put a reverse proxy in front of two machines and write rules for health checks, retries and stickiness. It works, and it becomes a small project you now maintain, plus you still need a way to reach both from outside your network.

One static endpoint, several machines

With a static endpoint, a named URL such as https://tunnel.tokmine.ai/t/acme-prod stays the same and several of your machines can share it. Start each machine against it:

tokmine --endpoint acme-prod

How requests are routed

  • Requests are routed by model name. The model must be served by at least one online machine.
  • When several machines serve it, the request goes to the one with the most headroom.
  • A machine that goes offline simply stops receiving traffic.

Clients see one URL and one token. Adding or removing a machine needs no client change.

Targeting a specific machine

Sometimes you want a particular provider, for example to test a new model on one box. Add /p/<providerId> after the slug to pin the request.

Try it

Static endpoints are for paid accounts. Setup steps are in the endpoints docs.