Tunnels are great tools. Exposing localhost:11434 with one command is how many people first use a model away from home. The trouble starts on day two.
Where the gaps show up
- Changing URLs. If the address changes on restart, every config that used it breaks.
- Session limits. A tunnel that closes after a while is a poor home for a model your editor calls all day.
- No auth. An open URL in front of a GPU is an invitation. Adding auth means adding another layer.
- No visibility. You cannot see who called the model or how much.
- No model awareness. The tunnel forwards bytes. It does not know which API dialect a client speaks or which machine serves which model.
What a model-aware layer adds
Tokmine finds the model servers on the machine, detects which API dialects they speak, and puts one URL in front of them that answers OpenAI, Anthropic and Gemini clients. It keeps the same URL when the machine reconnects, requires a token on every request, and shows usage on a dashboard.
When a plain tunnel is enough
If it is a quick demo, one person, one afternoon, a simple tunnel is fine. Anonymous Tokmine sessions cover that case too, for 30 minutes without an account.
Try it
Run tokmine and compare. The quickstart shows the whole flow.