Skip to content

Lightweight · High-performance · Low-cost

aiXapply-4B · Code-change application model

Released and open-sourced in May 2026. A 4B model whose accuracy rivals hundred-billion-parameter models, with 15× faster inference on a single consumer GPU.

aiXapply-4B is purpose-built for the high-frequency “code-change application” scenario in agent workflows: merging model-generated snippets precisely into the original file while fully preserving existing formatting and surrounding structure

BENCHMARKS

Benchmarks

Accuracy on par with hundred-billion-parameter models
aiXapply-4B-SFT reaches 94.4% accuracy — ahead of DeepSeek-V3.2-671B (92.5%) and GLM-5-744B (92.1%), and far ahead of the same-size Qwen3-4B (62.6%).
Inference dozens of times faster
On a single A100 40GB, inference runs at 2,692 tokens/sec and task latency drops from minutes to seconds — fast enough for real-time agent workflows.
Stable cross-scenario generalization
Remains stable on long-context adaptation and on programming languages that make up a tiny share of training data, matching the variety of real enterprise codebases.

TALK TO US

Efficient, intelligent R&D built on your own codebase.

Needs assessment → POC on your enterprise codebase → private rollout, deployed on your own compute (domestic chips included).