Thursday, August 27, 2026
Tech Beat
Aug 20, 2026, 4:52 PMArtificial Intelligence

Liquid AI DSpark Speeds LFM2.5 Inference Up to 3.18x

Liquid AI's DSpark checkpoints boost LFM2.5 inference up to 3.18x on H100 and 2.87x on M4 Max, while preserving identical greedy outputs across five tests.

Listen to this briefingAudio briefing

Summary

On August 20, 2026, Liquid AI researchers Leonie Monigatti, Fernando Fernandes Neto, Tarek Dakhran and Nathan Ranchin released DSpark checkpoints for LFM2.5-1.2B-Instruct, LFM2.5-2.6B and LFM2.5-8B-A1B. DSpark combines a DFlash-style parallel backbone, sequential Markov head and confidence-scheduled verifier to draft candidates for one target-model verification pass and prune uneconomic suffixes. Each attention-only drafter uses 5 layers, block size 9 and about 300M parameters, specifically 295.7M, 327.7M and 327.7M, after 15 epochs on SFT, chat, code and function-calling data. Checkpoints were selected by acceptance rate. Verification makes temperature-0 greedy output identical to the baseline, preserving pass@1 and exact-match accuracy for minimal added memory.

Liquid AI tested up to 256 output tokens at batch size 1 on MATH500, HumanEval, MBPP, GSM8K and MT-Bench, using SGLang in BF16 on one H100 80 GB and llama.cpp with experimental Metal kernels and FP16 GGUF on an M4 Max MacBook Pro. Mean accepted drafts out of 10 were 4.81, 5.02 and 6.95 for the 2.6B, 1.2B and 8B models. LFM2.5-2.6B averaged 2.67x on H100, 323 to 864 tok/s, and 2.27x on M4 Max, 61 to 139 tok/s, while cutting multi-tool function-calling latency 57%. LFM2.5-1.2B-Instruct averaged 2.10x and 2.54x, peaking at 2.87x on-device, although speedup varied 52%. LFM2.5-8B-A1B averaged 2.54x on H100 and 1.18x on M4 Max, but reached 3.18x on MATH500 on H100, 428 to 1,362 tok/s. Metal’s MoE implementation and extra expert traffic limited Mac gains.

Open-source support arrived through SGLang PR #31041 and llama.cpp PR #27383. Hugging Face offers all three checkpoints as Safetensors and GGUF, with an OpenAI-compatible SGLang endpoint.

Positives

  • LFM2.5-8B-A1B reached 3.18x H100 throughput on MATH500, increasing performance from 428 to 1,362 tok/s.
  • LFM2.5-1.2B-Instruct peaked at 2.87x on the M4 Max, rising from 136 to 389 tok/s on HumanEval.
  • LFM2.5-2.6B reduced average multi-tool function-calling latency by 57%, supporting more responsive on-device agents.
  • Target-model verification preserves identical greedy output, leaving pass@1 and exact-match benchmark accuracy unchanged.
  • SGLang PR #31041 and llama.cpp PR #27383 provide open-source DSpark support from launch day.
  • Safetensors and GGUF checkpoints are available on Hugging Face for all three LFM2.5 models.

Risks & concerns

  • LFM2.5-8B-A1B averaged only 1.18x faster on M4 Max despite averaging 2.54x on H100.
  • llama.cpp’s Metal MoE implementation and increased expert weight traffic restricted on-device gains for LFM2.5-8B-A1B.
  • LFM2.5-1.2B-Instruct speedup varied by as much as 52% because acceptance rates changed with the underlying text distribution.
  • Published benchmarks used only batch size 1, temperature 0 and up to 256 output tokens.
  • Each speculative path adds a 295.7M or 327.7M parameter draft model and a small memory overhead.
  • The llama.cpp measurements relied on experimental Metal kernels.
Primary sourceHugging Face - Bloghttps://huggingface.co/blog/LiquidAI/lfm25-dspark
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceAug 27

OpenAI Brings ChatGPT Ads to India With 50 Brands, ₹725 Daily Floor

Artificial IntelligenceAug 27

Nvidia Nears $12.9 Billion Hugging Face Acquisition Amid Conflicting Reports

Artificial IntelligenceAug 27

OpenAI Expands Brazil Presence to Support Nationwide AI Adoption