Thursday, August 27, 2026
Tech Beat
Aug 19, 2026, 1:48 PMArtificial Intelligence

Liquid AI LFM2.5 QAD Q4_0 Models Recover 97% of BF16 Accuracy

Liquid AI releases QAD Q4_0 checkpoints for four LFM2.5 models, recovering about 97% of BF16 accuracy while preserving 4-bit edge speed and model size.

Listen to this briefingAudio briefing

Summary

Liquid AI authors Aditya Tadimeti and Leonie Monigatti released Quantization-Aware Distillation, or QAD, Q4_0 GGUF checkpoints on August 19, 2026, for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B. QAD distills a high precision teacher into a quantized student, preserving native Q4_0 memory use and speed while recovering 97% of average accuracy lost through quantization. Against released post-training quantization GGUFs, the four checkpoints retained 97.1%, 96.5%, 97.4%, and 96.6% of their respective BF16 baselines, based on five-run means across GPQA Diamond, MMLU-Pro, IFEval, IFBench, Multi-IF, BFCLv4, plus GSM8K for the 230M and 350M models and AIME25 for the 1.2B and 2.6B models. BF16 GGUF was the in-format ceiling.

Tests covered GPU inference on MacBook Pro and NucBox EVO-X2, plus Arm CPU inference on Samsung Galaxy S26 Ultra and Raspberry Pi 5, with BF16 and F16 references where profiled. The 230M and 350M QAD Q4_0 models matched Q5_K_M quality within evaluation variance while decoding 4% to 33% faster; the 1.2B and 2.6B models matched Q4_K_M quality with 3% to 14% higher throughput. Applicable 230M and 1.2B checkpoints also matched Unsloth UD-Q4_K_XL. All four files are available on Hugging Face for llama.cpp and other runtimes supporting GGUF Q4_0 artifacts.

Positives

  • 97.1%, 96.5%, 97.4%, and 96.6% of BF16 performance was retained by the four respective LFM2.5 QAD checkpoints.
  • 4% to 33% higher decode throughput accompanied Q5_K_M level quality for the 230M and 350M models.
  • 3% to 14% higher throughput accompanied Q4_K_M level quality for the 1.2B and 2.6B models.
  • Four hardware targets covered Apple and x86 GPU inference plus Samsung and Raspberry Pi Arm CPU inference.
  • All four QAD Q4_0 GGUF files became available on Hugging Face for llama.cpp and compatible runtimes.

Risks & concerns

  • 96.5% BF16 retention for LFM2.5-350M was the weakest result among the four QAD checkpoints.
  • Unsloth UD-Q4_K_XL comparisons applied only to the 230M and 1.2B models.
  • BF16 and F16 references were included only where profiled, limiting full precision comparisons across every hardware target.
  • The results came from Liquid AI’s own blog, and the article described no independent validation.
Primary sourceHugging Face - Bloghttps://huggingface.co/blog/LiquidAI/qad
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceAug 27

OpenAI Brings ChatGPT Ads to India With 50 Brands, ₹725 Daily Floor

Artificial IntelligenceAug 27

Nvidia Nears $12.9 Billion Hugging Face Acquisition Amid Conflicting Reports

Artificial IntelligenceAug 27

OpenAI Expands Brazil Presence to Support Nationwide AI Adoption