Thursday, August 27, 2026
Tech Beat
Aug 25, 2026, 11:39 AMArtificial Intelligence

Multiverse QAH Makes a 4-Bit GPT-OSS Model Beat Its Full-Precision Source

Multiverse Computing's QAH turns a compressed GPT-OSS 120B into a 60B MXFP4 model that beats its BF16 checkpoint on seven of nine benchmarks at lower cost.

Listen to this briefingAudio briefing

Summary

On August 25, 2026, Multiverse Computing researchers Antonio Tiene, Iker García-Ferrero, Ali Hashemi and Bakbergen Ryskulov introduced Quantization-Aware Healing, or QAH. They compressed GPT-OSS 120B to 60B parameters, recovered it in bfloat16, then quantized it to MXFP4 while distilling KL-divergence logits directly from the frozen original model. Unlike QAT, QAH uses no hard labels or repeated supervised fine-tuning, RLHF or agentic tuning. Unlike conventional quantization-aware distillation, it avoids using the recovered checkpoint as teacher. A chunked KL loss processes documents up to 32k tokens without materializing the full vocabulary-by-sequence grid.

The 60B MXFP4 model beat its BF16 source on seven of nine benchmarks: AA-LCR, 42.7 versus 35.3, up 7.4; AIME 2025, 76.3 versus 70.7, up 5.6; Aider, 40.9 versus 38.2, up 2.7; τ²-bench, 61.7 versus 59.4, up 2.3; GPQA Diamond, 67.4 versus 65.7, up 1.7; IFBench, 59.9 versus 58.4, up 1.5; and LiveCodeBench, 66.5 versus 65.5, up 1.0. It trailed on MMLU-Pro, 73.8 versus 74.0, and SciCode, 34.2 versus 35.6. It also beat the 120B teacher on LiveCodeBench, 66.5 versus 66.0, and approached its 69.0 GPQA score. On GPT-OSS 9B, QAH peaked at 54.9 after about 100 steps and remained within two points, while QAT reached 54.6 near 700 steps and lost almost 19 points by step 1,200. QAH uses roughly four times less weight memory than BF16 and half the teacher's compute per token; combined reductions could approach eight times for BF16 model families.

Positives

  • Seven of nine benchmarks favored the 60B MXFP4 QAH model over its recovered BF16 checkpoint.
  • AA-LCR improved 7.4 points and AIME 2025 gained 5.6 points, recovering capabilities commonly damaged by compression.
  • LiveCodeBench reached 66.5, beating both the 60B BF16 model at 65.5 and the 120B teacher at 66.0.
  • QAH peaked at 54.9 after about 100 steps, roughly seven times sooner than QAT reached 54.6.
  • Four times less weight memory than BF16 and half the 120B teacher's compute per token enable substantially smaller deployment hardware.

Risks & concerns

  • MMLU-Pro fell 0.2 points against the BF16 checkpoint, while SciCode declined 1.4 points.
  • AA-LCR remained 7.3 points behind the 120B teacher, showing that extreme long-context capacity remains difficult to recover.
  • QAT lost nearly 19 points by step 1,200, creating a deployment risk without careful validation and early stopping.
  • Results presented cover GPT-OSS 120B and 9B experiments, leaving performance on other architectures and workloads untested.
Primary sourceHugging Face - Bloghttps://huggingface.co/blog/MultiverseComputingCAI/quantization-aware-healing
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceAug 27

OpenAI Brings ChatGPT Ads to India With 50 Brands, ₹725 Daily Floor

Artificial IntelligenceAug 27

Nvidia Nears $12.9 Billion Hugging Face Acquisition Amid Conflicting Reports

Artificial IntelligenceAug 27

OpenAI Expands Brazil Presence to Support Nationwide AI Adoption