XCurOS 2.2Max

Intelligence that never leaves your network.

A secure multimodal model for enterprise and on-premise deployment. Advanced reasoning, native vision and document understanding, agentic tool use, and a context of up to 1M tokens — running entirely behind your own firewall.

Fully isolated deployment OpenAI-compatible API Vision + documents + video

Works with the coding agents you already use

010BTotal parameters, Mixture-of-Experts
02~0BActive per token — ~9× less compute
031MMaximum context the model takes — we serve it at 200K
040Experts, top-8 routed per token
The family01

The newest model, and the one it grew from.

XCurOS 2.2 MaxLatest · what we serve
XCurOS-2.2-A4B-Max

Post-trained from A4B through a self-improvement loop — the strongest agentic coding, tool use and long-horizon planning in the lineup. Ahead on 52 of the 56 published benchmarks.

XCurOS 2.2 A4BBase tier
XCurOS-2.2-A4B

The 2.2 checkpoint with the general-purpose post-training: reasoning, native vision, documents and OCR, tool calling, and the full context window.

Compare the two tiers

Capabilities02

One model for the work your business actually does.

Text, images, documents and video in the same conversation — with the tool calling and long context that real workflows need.

01

Security-first design

Built to run fully isolated, on-premise, with complete data privacy — nothing leaves your infrastructure.

02

Advanced reasoning

Competitive with the strongest open-weight models on mathematics, science and complex multi-step reasoning.

03

Native vision

Reads images, screenshots, charts, tables and long documents, with strong OCR and video comprehension.

04

Agentic & tool-ready

Reliable function calling, MCP, repository-level coding and OS-level automation workflows.

05

Extreme long context

Up to 1,048,576 tokens — book-length references and whole codebases in one conversation. The endpoint we run is set to 200K.

06

Efficient MoE

~4B of ~35B parameters activate per token — roughly 9× less compute than a dense model of the same size.

07

Multilingual

Robust understanding and generation across a wide range of languages, including Arabic and English.

08

Drop-in integration

OpenAI-compatible API — slot it into the pipelines you already run, with no code changes.

Measured, not claimed03

Every benchmark. Every model. The protocol attached.

Every XCurOS figure but one was measured in-house at BF16 on 2× A100 80 GB — AA-LCR is provider-run, on the publisher's own infrastructure. Competitor scores are transcribed from their publishers, so read the tables as capability comparisons rather than controlled experiments — the full protocol is on the benchmarks page.

XCurOS2.2 Max is the highlighted bar, with the A4B tier charted beside it · higher is better · -- means the score was not published Full model card
Deployment04

Your weights. Your hardware. Your data — start to finish.

XCurOS is designed for the environments where sending a prompt to a third party is simply not an option: regulated industries, classified networks, and air-gapped estates.

01

Runs on your own GPUs

From 2× 80 GB at BF16 down to two 16 GB cards with the Q4_K_M build — which is what runs the public endpoint. No egress, no shared tenancy.

02

No telemetry, no retention

The serving stack writes nothing back to us. Conversations in the playground stay in the browser that made them.

03

Predictable economics

A sparse model with a small KV cache: ~20 KB per token against ~80 KB if all 40 layers held one.

Integration05

Three lines to switch.

The endpoint speaks the OpenAI protocol, and the native Anthropic Messages API too. Point your existing SDK at it and keep the rest of your code exactly as it is.


        
System: contact

Run it inside your network.

Open the playground and try it right now — it runs at a 200K context — or talk to us about a licence, an evaluation, or an on-premise deployment.