Docs

Install a local model

Tokmine exposes a model server that already runs on your machine. If you do not have one yet, tokmine install sets one up for you: it installs Ollama, picks a model that fits your hardware and downloads it.

The short version

bash
tokmine install ollama   # install Ollama and pull a model that fits this machine
tokmine install glm      # or pick a model family
tokmine                  # expose it

If you run tokmine and no local model server is found (and you passed no --provider), the CLI prints No local model server found. Install one in a minute: tokmine install ollama. In an interactive terminal it then asks Install Ollama now? [y/N] and, if you agree, shows the full plan and asks a second time before doing anything. This is skipped with --json, --no-discover or when input is not a terminal. If a server is running but has no models, it suggests tokmine install --model <name>.

What it does

  1. Detects a runner that is already running or installed and skips straight to the model download.
  2. Installs the runner. The default is Ollama. The CLI picks the best method for your system (see below).
  3. Picks a model from the machine's RAM and GPU memory, unless you name one with --model, and shows it with the download size.
  4. Pulls the model with progress. This can be several gigabytes. Ctrl-C cancels, and running it again resumes the download.
  5. Then you run tokmine, which discovers the new server like any other.
Nothing runs silently. The CLI prints every command and URL it will use (official HTTPS sources only), notes when a script uses sudo, and asks you to confirm with y. Pass --yes to skip the prompt. Without a terminal, --yes is required and the CLI refuses otherwise. --dry-run only prints the plan.

Options

text
tokmine install [runner|family] [--model NAME] [--method auto|script|brew|winget|docker] [--yes] [--dry-run]
OptionDescription
[runner]ollama (default) or llamacpp (Docker only). With no argument in a non-interactive shell, it prints the options and exits with code 2, unless --yes, --dry-run, --model or --method is set.
[family]A model family installs Ollama if missing and pulls a model from that family that fits the machine: glm (also glm4), qwen, llama, gemma, mistral, phi, deepseek, gpt-oss. GLM has no small tag, so small machines fall back to glm4:9b.
--model NAMEPull this model instead of choosing one from your hardware, for example llama3.2:3b.
--methodauto (default), script, brew, winget or docker.
--yesDo not ask for confirmation.
--dry-runPrint the plan and change nothing.

Which model it picks

MachineModel
Less than 8 GB RAMllama3.2:3b
8 to 16 GB RAMllama3.1:8b
About 12 GB VRAM, 24 GB unified memory or 32 GB RAMqwen3:14b
About 24 GB VRAM, 48 GB unified memory or 64 GB RAMqwen3:32b

Steps per operating system

With --method auto the CLI tries these in order and shows what it will run.

Linux

Ollama's official install script (curl -fsSL https://ollama.com/install.sh | sh), which uses sudo and asks for your password itself. If that is not possible, Docker.

macOS

brew install ollama plus brew services start ollama. Without Homebrew, the official app zip is unpacked into ~/Applications, which needs no admin rights. Otherwise Docker.

Windows

winget install --id Ollama.Ollama -e. Otherwise Ollama's official PowerShell installer (irm https://ollama.com/install.ps1 | iex), then Docker.

Docker

--method docker runs the runner in a container, and is the last fallback for Ollama. The container is bound to 127.0.0.1 only, so it is not exposed on your network, and tokmine discovers it as usual. --gpus=all is added when Docker has the NVIDIA runtime (Linux and Windows). llamacpp is always installed this way, using the ghcr.io/ggml-org/llama.cpp:server image (the -cuda variant when the NVIDIA runtime is present), which loads the model from Hugging Face on first start.

bash
tokmine install ollama --method docker
tokmine install llamacpp
tokmine install ollama --dry-run     # preview only

GPU notes

  • NVIDIA: install the current driver first. Ollama on Linux and Windows uses it automatically. For Docker you also need the NVIDIA Container Toolkit so the CLI can pass --gpus=all.
  • AMD: supported by Ollama on Linux and Windows with recent drivers. Support depends on the card.
  • Apple Silicon: the GPU is used automatically, and model size is chosen from unified memory. Docker on macOS cannot use it, so prefer the default methods.
  • No GPU: everything works on the CPU with a smaller model and slower replies.

Manual alternatives

You do not need tokmine install. Install any runner yourself, start it, and run tokmine. Discovery finds it on its default port (see Supported runners).

RunnerGet started
OllamaInstall from ollama.com/download, then ollama pull llama3.2.
LM StudioA desktop app, so tokmine install cannot install it. Download from lmstudio.ai, load a model and turn on the local server in the Developer tab.
llama.cppBuild or download llama-server and run llama-server -m model.gguf --port 8080.
vLLMpip install vllm, then vllm serve <model>. Needs a supported GPU.

Troubleshooting

  • tokmine install says the runner is installed but tokmine finds no models.
    The install may have pulled no model yet, or the server is not running. Run tokmine models to see what is discovered, and start the runner (for example ollama serve) if needed.
  • The model download is slow or fails.
    Models are several gigabytes. Check your connection and free disk space, then run the same command again, which resumes the download. Behind a proxy, set HTTPS_PROXY.
  • It chose a model that is too big or too small.
    Pass --model NAME to choose one yourself. Use --dry-run to see what would be picked for this machine. GLM has no small tag, so small machines get glm4:9b.
  • Permission denied or sudo prompts on Linux.
    The official Linux script uses sudo and asks for your password itself. Run it from a terminal where you can enter it, or use --method docker.
  • --method docker fails.
    Docker must be installed and its daemon running, and your user must be allowed to use it. On Linux with an NVIDIA GPU, install the NVIDIA Container Toolkit.
  • It refuses to run and asks for --yes.
    Without a terminal the CLI cannot ask for confirmation. Add --yes to accept the plan, or --dry-run to only print it.
  • I want to review the script before it runs.
    The CLI always shows it and asks before running. Answer no, then run it yourself, or use one of the manual alternatives.

When it works, continue with the Quickstart, then try your endpoint in the browser.