Skip to main content
The model is the brain doing the thinking inside your framework. You pick one at deploy time and can change it later.
Agents need tool-calling. A model that can’t call tools can’t run the agent loop, so the picker only offers models verified to work end to end.

The lineup

Inference is routed through Voight’s own key — you don’t bring one. Two families are available:

GLM — by Z.ai

The GLM lineup (4.5, 4.6, 4.7, 5.x). Strong general reasoning and tool use. Sponsored by Z.ai.

MiMo — by Xiaomi

Fast, inexpensive, tool-capable, with a very large context window.
The default is GLM 4.6: balanced, proven in production, and strong at tool use. MiMo stays available as the cheapest per-token option. Every new account starts with a welcome inference credit, and usage past it draws from your prepaid credit balance: you always see what a model costs before you pick it.

Bring your own key — coming soon

The next phase opens the model choice wide open, with cost billed to your own provider account:
  • BYOK · OpenRouter — paste your OpenRouter key and pick any model in the OpenRouter catalog.
  • BYOK · Direct — connect a provider key directly: OpenAI, Anthropic, Gemini, Qwen, Z.ai, or Xiaomi. Calls go straight to the vendor.
Per-model cost controls land alongside token metering in the same phase.

Choosing

Next