Agents need tool-calling. A model that can’t call tools can’t run the agent loop, so the picker only offers models verified to work end to end.
The lineup
Inference is routed through Voight’s own key — you don’t bring one. Two families are available:GLM — by Z.ai
The GLM lineup, 4.5 through 5.3 and 5.3 Flash. Strong general reasoning and tool use. Sponsored by Z.ai.
MiMo — by Xiaomi
Fast, inexpensive, tool-capable, with a very large context window.
On a Nosana GPU: local, private
Agents hosted on a Nosana GPU don’t use the sponsored lineup. Hermes agents run Hermes 3 8B (NousResearch’s Llama 3.1 8B fine-tune, trained for the Hermes runtime’s tool-calling format) and ZeroClaw agents run Qwen2.5 7B Instruct, in both cases locally on the job’s own GPU. Inference never leaves the GPU and no provider key exists in the container. It is the private, verifiable option; for heavy reasoning, the GLM lineup on Voight Cloud is stronger.Bring your own key
Open the model choice wide, with inference billed to your own provider account. Both modes live in the agent’s Settings:- BYOK · OpenRouter — paste your OpenRouter key and pick any model in the OpenRouter catalog, GLM 5.3 and 5.3 Flash included.
- BYOK · Direct — connect a provider key directly: OpenAI, Anthropic, Gemini, Qwen, Z.ai, or Xiaomi. Calls go straight to the vendor (GLM 5.3 and 5.3 Flash via your Z.ai key).
Choosing
Next
- Frameworks — the engine the model runs inside
- Channels — where the agent works
- Pricing — what a running agent costs