Lightweight · High-performance · Low-cost
aiXapply-4B · Code-change application model
Released and open-sourced in May 2026. A 4B model whose accuracy rivals hundred-billion-parameter models, with 15× faster inference on a single consumer GPU.
aiXapply-4B is purpose-built for the high-frequency “code-change application” scenario in agent workflows: merging model-generated snippets precisely into the original file while fully preserving existing formatting and surrounding structure
BENCHMARKS
Benchmarks
- Accuracy on par with hundred-billion-parameter models
- aiXapply-4B-SFT reaches 94.4% accuracy — ahead of DeepSeek-V3.2-671B (92.5%) and GLM-5-744B (92.1%), and far ahead of the same-size Qwen3-4B (62.6%).
- Inference dozens of times faster
- On a single A100 40GB, inference runs at 2,692 tokens/sec and task latency drops from minutes to seconds — fast enough for real-time agent workflows.
- Stable cross-scenario generalization
- Remains stable on long-context adaptation and on programming languages that make up a tiny share of training data, matching the variety of real enterprise codebases.
