Models

Open-weight frontier models, served from our iron.

Every model on Zarx has published weights and a license you can read. These are the models we run today; the launch catalog for design partners is confirmed with each partner.

Catalog

Models we run today include

ModelVendorFamilyContextLicenseInputServed with
GLM-5.2glm-5.2Z.aiMixture of experts, long-horizon and coding1M tokensMITTextvLLM (FP8 build)
Kimi K2.6kimi-k2.6Moonshot AIMixture of experts, 1T total / 32B active, agentic256K tokensModified MITText, imageSGLang
Gemma 4 12Bgemma-4-12bGoogle DeepMindDense, multimodal, multilingual256K tokensApache 2.0Text, image, audiovLLM (FP8 build)
  • GLM-5.2: GLM-5.2 (Z.ai): mixture-of-experts model for long-horizon and coding tasks, MIT license, 1M-token context. We run the FP8 build on vLLM. Model card
  • Kimi K2.6: Kimi K2.6 (Moonshot AI): mixture-of-experts model with 1T total and 32B active parameters, text and image input, Modified MIT license, 256K-token context. We run it on SGLang. Model card
  • Gemma 4 12B: Gemma 4 (Google DeepMind): open-weight family with text, image and audio input, Apache 2.0 license; the 12B build we have run on vLLM has a 256K-token context. Model card

How models land

Four rules for the catalog.

  • Open weights only

    We serve models whose weights are published under a license we can read. That is what makes the jurisdiction claim real: there is no remote model API behind Zarx.

  • Pinned checkpoints

    A model id maps to one published checkpoint. New weights get a new id. Your evaluations stay valid until you decide to move.

  • We run what we can stand behind

    Before a model reaches the catalog it runs on our iron with our serving stack, under load, for our own products. The catalog grows at the pace we can operate, not at the pace of announcements.

  • Fine-tunes stay where the base model runs

    Design partners who fine-tune do so on the same servers that serve the base model. Training data, adapters and outputs never leave NPAW iron.

Serving stack: Models are served with vLLM and SGLang, open-source inference engines, on ROCm for the MI300X. Source: NPAW infrastructure inventory and model registry (vLLM and SGLang on MI300X hosts).

Requesting a model

Need a model that is not listed?

Design partners can ask for any open-weight model whose license permits hosted serving. We evaluate it on our iron, with our serving stack, and tell you what we can operate. Models we cannot stand behind in production do not enter the catalog, however loud the release.

Fine-tunes of catalog models are served under a partner-specific id, on the same servers as the base model. Adapters and training data stay on NPAW iron.

Tell us which models you need in your jurisdiction.

Request access and name the models and workloads you run today. We reply with what we can serve and when.