Models
Open-weight frontier models, served from our iron.
Every model on Zarx has published weights and a license you can read. These are the models we run today; the launch catalog for design partners is confirmed with each partner.
Catalog
Models we run today include
| Model | Vendor | Family | Context | License | Input | Served with |
|---|---|---|---|---|---|---|
GLM-5.2glm-5.2 | Z.ai | Mixture of experts, long-horizon and coding | 1M tokens | MIT | Text | vLLM (FP8 build) |
Kimi K2.6kimi-k2.6 | Moonshot AI | Mixture of experts, 1T total / 32B active, agentic | 256K tokens | Modified MIT | Text, image | SGLang |
Gemma 4 12Bgemma-4-12b | Google DeepMind | Dense, multimodal, multilingual | 256K tokens | Apache 2.0 | Text, image, audio | vLLM (FP8 build) |
- GLM-5.2: GLM-5.2 (Z.ai): mixture-of-experts model for long-horizon and coding tasks, MIT license, 1M-token context. We run the FP8 build on vLLM. Model card
- Kimi K2.6: Kimi K2.6 (Moonshot AI): mixture-of-experts model with 1T total and 32B active parameters, text and image input, Modified MIT license, 256K-token context. We run it on SGLang. Model card
- Gemma 4 12B: Gemma 4 (Google DeepMind): open-weight family with text, image and audio input, Apache 2.0 license; the 12B build we have run on vLLM has a 256K-token context. Model card
How models land
Four rules for the catalog.
Open weights only
We serve models whose weights are published under a license we can read. That is what makes the jurisdiction claim real: there is no remote model API behind Zarx.
Pinned checkpoints
A model id maps to one published checkpoint. New weights get a new id. Your evaluations stay valid until you decide to move.
We run what we can stand behind
Before a model reaches the catalog it runs on our iron with our serving stack, under load, for our own products. The catalog grows at the pace we can operate, not at the pace of announcements.
Fine-tunes stay where the base model runs
Design partners who fine-tune do so on the same servers that serve the base model. Training data, adapters and outputs never leave NPAW iron.
Serving stack: Models are served with vLLM and SGLang, open-source inference engines, on ROCm for the MI300X. Source: NPAW infrastructure inventory and model registry (vLLM and SGLang on MI300X hosts).
Requesting a model
Need a model that is not listed?
Design partners can ask for any open-weight model whose license permits hosted serving. We evaluate it on our iron, with our serving stack, and tell you what we can operate. Models we cannot stand behind in production do not enter the catalog, however loud the release.
Fine-tunes of catalog models are served under a partner-specific id, on the same servers as the base model. Adapters and training data stay on NPAW iron.
Tell us which models you need in your jurisdiction.
Request access and name the models and workloads you run today. We reply with what we can serve and when.