Thursday, September 24, 2026
Tech Beat
Sep 24, 2026, 2:08 PMArtificial Intelligence

Liquid AI LFM2.5-VL-DSpark Speeds Vision AI Decoding Up to 3.13x

Liquid AI's 280M-parameter LFM2.5-VL-DSpark accelerates LFM2.5-VL-3B decoding up to 3.13x on devices and 2.66x on H100 with exact output across three runtimes.

Listen to this briefingAudio briefing

Summary

On September 24, 2026, Liquid AI researchers Yuri Khrustalev, Leonie Monigatti and Viviana Márquez released experimental LFM2.5-VL-3B-DSpark for the 3B LFM2.5-VL-3B vision-language model. The attention-only drafter taps target hidden states after image patches and text tokens enter a shared representation, then proposes token blocks that the target verifies, preserving greedy output. Trained for 10 epochs on weighted vision-language SFT data after 3-layer, 4-layer and 5-layer ablations, it uses 4 layers and block size 9, with 8 or 9 recommended at inference.

Its 279.5M parameters, rounded to 280M or 8.9% overhead, comprise a 193.0M decoder, 21.0M hidden-state projection, 65.5M Markov head and 6.4k norms plus confidence head. With block size 8 across six MMSpec tasks, general VQA, text VQA, captioning, chart VQA, complex reasoning and multi-turn conversation, MLX on M5 Max delivered 2.30x to 3.13x faster decoding and 1.56x to 2.62x end-to-end gains. llama.cpp on M3 Ultra reached 1.57x to 2.14x and 1.30x to 1.77x, respectively.

On H100, SGLang produced up to 2.66x decode and 1.64x to 2.27x end-to-end gains, although the detailed decode range is printed as 20.4x to 2.66x, conflicting with the stated maximum. DSpark cannot accelerate vision encoding or prefill, which consume more latency on edge hardware, though M5 per-core GPU neural accelerators narrow that gap. Integrations require SGLang PR #40651, llama.cpp PR #29339 or MLX-VLM PR #2280. The open-weight model is available on Hugging Face in Safetensors and GGUF for unrestricted downloading, fine-tuning and deployment.

Positives

  • M5 Max decoding improved 2.30x to 3.13x, while complete request latency improved 1.56x to 2.62x.
  • H100 testing delivered up to 2.66x faster decoding and 2.27x end-to-end acceleration.
  • 279.5M added parameters increase the 3B target model’s deployed parameter count by only 8.9%.
  • Target verification makes speculative decoding exact, so greedy output matches LFM2.5-VL-3B running alone.
  • llama.cpp, MLX-VLM and SGLang integrations arrived alongside Safetensors and GGUF files on Hugging Face.

Risks & concerns

  • Vision encoding and prefill remain unaccelerated, limiting end-to-end improvements despite larger decoding gains.
  • Edge devices spend more total latency on prefill than datacenter GPUs because they offer substantially less compute.
  • H100 results list a 20.4x to 2.66x decode range that conflicts with the separately stated 2.66x maximum.
  • SGLang, llama.cpp and MLX-VLM each require specific builds or pull requests rather than standard releases.
  • The experimental drafter adds 279.5M parameters and therefore still raises deployment memory requirements by 8.9%.
Primary sourceHugging Face - Bloghttps://huggingface.co/blog/LiquidAI/lfm2-5-vl-dspark
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceSep 24

Gemini 3.8 Live Avatar Brings 97-Language AI Video to Enterprise

Artificial IntelligenceSep 24

Google Project Suncatcher Satellite to Test Orbital AI on October 1

Artificial IntelligenceSep 24

Shield AI, Waabi and GM to Detail Safety Tests for Physical AI at Disrupt 2026