One GPU box eventually gets busy. The usual fix, a second box, creates a new question: which one answers?
The DIY version
You can put a reverse proxy in front of two machines and write rules for health checks, retries and stickiness. It works, and it becomes a small project you now maintain, plus you still need a way to reach both from outside your network.
One static endpoint, several machines
With a static endpoint, a named URL such as https://tunnel.tokmine.ai/t/acme-prod stays the same and several of your machines can share it. Start each machine against it:
tokmine --endpoint acme-prod
How requests are routed
- Requests are routed by model name. The model must be served by at least one online machine.
- When several machines serve it, the request goes to the one with the most headroom.
- A machine that goes offline simply stops receiving traffic.
Clients see one URL and one token. Adding or removing a machine needs no client change.
Targeting a specific machine
Sometimes you want a particular provider, for example to test a new model on one box. Add /p/<providerId> after the slug to pin the request.
Try it
Static endpoints are for paid accounts. Setup steps are in the endpoints docs.