You have Ollama on a machine with a good GPU. Here is how to use it from anywhere.

1. Install the CLI

curl -fsSL https://tokmine.ai/install.sh | sh

On Windows, use irm https://tokmine.ai/install.ps1 | iex.

2. Run it

tokmine

The CLI finds Ollama on port 11434, detects that it speaks both the Ollama and OpenAI dialects, and prints a public URL such as https://tunnel.tokmine.ai/t/<slug> with a token. Without an account this lasts 30 minutes.

3. Point your client at it

from openai import OpenAI

client = OpenAI(
    base_url="https://tunnel.tokmine.ai/t/<slug>/v1",
    api_key="<token>",
)
client.chat.completions.create(
    model="llama3.2",
    messages=[{"role": "user", "content": "Hello"}],
)

Keeping it

Create an account and run tokmine login to keep the machine online permanently. It costs $0.50 per machine per month, plus $0.50 per person who calls it. Read more in the API compatibility docs.