Skip to main content
The model is the brain doing the thinking inside your framework. You pick one at deploy time and can change it later.
Agents need tool-calling. A model that can’t call tools can’t run the agent loop, so the picker only offers models verified to work end to end.

The lineup

Inference is routed through Voight’s own key — you don’t bring one. Two families are available:

GLM — by Z.ai

The GLM lineup, 4.5 through 5.3 and 5.3 Flash. Strong general reasoning and tool use. Sponsored by Z.ai.

MiMo — by Xiaomi

Fast, inexpensive, tool-capable, with a very large context window.
The default is GLM 4.6: balanced, proven in production, and strong at tool use. MiMo stays available as a low-cost option. New: GLM 5.3 and GLM 5.3 Flash. The newest Z.ai flagship and its fast variant, both with a 1M-token context window and full tool-calling. GLM 5.3 is the strongest reasoner in the lineup; 5.3 Flash is fast and inexpensive. Both are available everywhere a model can be picked: on Voight’s key, on your own OpenRouter key, or on your own Z.ai key. Every new account starts with a welcome inference credit, and usage past it draws from your prepaid credit balance: you always see what a model costs before you pick it.

On a Nosana GPU: local, private

Agents hosted on a Nosana GPU don’t use the sponsored lineup. Hermes agents run Hermes 3 8B (NousResearch’s Llama 3.1 8B fine-tune, trained for the Hermes runtime’s tool-calling format) and ZeroClaw agents run Qwen2.5 7B Instruct, in both cases locally on the job’s own GPU. Inference never leaves the GPU and no provider key exists in the container. It is the private, verifiable option; for heavy reasoning, the GLM lineup on Voight Cloud is stronger.

Bring your own key

Open the model choice wide, with inference billed to your own provider account. Both modes live in the agent’s Settings:
  • BYOK · OpenRouter — paste your OpenRouter key and pick any model in the OpenRouter catalog, GLM 5.3 and 5.3 Flash included.
  • BYOK · Direct — connect a provider key directly: OpenAI, Anthropic, Gemini, Qwen, Z.ai, or Xiaomi. Calls go straight to the vendor (GLM 5.3 and 5.3 Flash via your Z.ai key).

Choosing

Next