Liquid AI LFM2.5 QAD Q4_0 Models Recover 97% of BF16 Accuracy
Liquid AI releases QAD Q4_0 checkpoints for four LFM2.5 models, recovering about 97% of BF16 accuracy while preserving 4-bit edge speed and model size.
Summary
Liquid AI authors Aditya Tadimeti and Leonie Monigatti released Quantization-Aware Distillation, or QAD, Q4_0 GGUF checkpoints on August 19, 2026, for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B. QAD distills a high precision teacher into a quantized student, preserving native Q4_0 memory use and speed while recovering 97% of average accuracy lost through quantization. Against released post-training quantization GGUFs, the four checkpoints retained 97.1%, 96.5%, 97.4%, and 96.6% of their respective BF16 baselines, based on five-run means across GPQA Diamond, MMLU-Pro, IFEval, IFBench, Multi-IF, BFCLv4, plus GSM8K for the 230M and 350M models and AIME25 for the 1.2B and 2.6B models. BF16 GGUF was the in-format ceiling.
Tests covered GPU inference on MacBook Pro and NucBox EVO-X2, plus Arm CPU inference on Samsung Galaxy S26 Ultra and Raspberry Pi 5, with BF16 and F16 references where profiled. The 230M and 350M QAD Q4_0 models matched Q5_K_M quality within evaluation variance while decoding 4% to 33% faster; the 1.2B and 2.6B models matched Q4_K_M quality with 3% to 14% higher throughput. Applicable 230M and 1.2B checkpoints also matched Unsloth UD-Q4_K_XL. All four files are available on Hugging Face for llama.cpp and other runtimes supporting GGUF Q4_0 artifacts.
Positives
- 97.1%, 96.5%, 97.4%, and 96.6% of BF16 performance was retained by the four respective LFM2.5 QAD checkpoints.
- 4% to 33% higher decode throughput accompanied Q5_K_M level quality for the 230M and 350M models.
- 3% to 14% higher throughput accompanied Q4_K_M level quality for the 1.2B and 2.6B models.
- Four hardware targets covered Apple and x86 GPU inference plus Samsung and Raspberry Pi Arm CPU inference.
- All four QAD Q4_0 GGUF files became available on Hugging Face for llama.cpp and compatible runtimes.
Risks & concerns
- 96.5% BF16 retention for LFM2.5-350M was the weakest result among the four QAD checkpoints.
- Unsloth UD-Q4_K_XL comparisons applied only to the 230M and 1.2B models.
- BF16 and F16 references were included only where profiled, limiting full precision comparisons across every hardware target.
- The results came from Liquid AI’s own blog, and the article described no independent validation.