Liquid AI’s 350M Model Gains 7.1 Points on IFStruct in 100 GRPO Steps
Liquid AI’s 350M model rises from 22.6% to 29.7% on IFStruct after 100 GRPO steps with 500 samples on a free 16 GB GPU, nearing Qwen3.5-2B’s 33.15%.
Summary
Published September 3, 2026, Leonie Monigatti, Ben Burtenshaw and Sergio Paniego present a public GitHub recipe that tunes Liquid AI’s LFM2.5-350M with TRL’s Group Relative Policy Optimization. IFStruct v1.0, available through Liquid4All/ifstruct and LiquidAI/ifstruct-v1.0, tests parseability and schema compliance. On an identical llama.cpp BF16 stack, the base model scored 452 of 2,000, or 22.6%, close to the published 21.1%, with 1,453 ms average latency. JSON reached 18.0%, YAML 27.2%, wrapper-key outputs 28.5% and bare lists 16.6%. Evaluation ran on an Apple M5 Max MacBook Pro with 36 GB unified memory.
Training used about 500 NVIDIA Nemotron structured-output samples. Forty percent added fenced-code instructions, while a separate 20% became top-level-array tasks. A rank-16 LoRA adapter trained about 6 million parameters, or 1.66% of the model, across 100 steps on a free-tier Colab or Kaggle 16 GB GPU. Configuration included a 5e-5 learning rate, 10 warmup steps, eight generations per group, batch size four, eight-step gradient accumulation, 1,024-token completions, temperature 1.1 and beta 0.01. Parseability, field-count and JSON Schema rewards were weighted 1.0, 0.5 and 2.0.
The merged BF16 model passed 594 of 2,000 tests, or 29.7%, a 7.1-point gain, at 1,518 ms latency. JSON rose 13.9 points to 31.9%, bare lists 13.1 points to 29.7%, wrapper keys 1.2 points to 29.7% and YAML 0.3 points to 27.5%. It remains below Qwen3.5-2B’s 33.15%. Required-field failures increased from 7,228 to 7,331, wrong item counts from 738 to 890 and type mismatches from 540 to 555, although bare-list wrapper errors fell from 100 to 102 in differently categorized totals and code-block errors dropped sharply.
Positives
- 594 of 2,000 IFStruct tests passed after tuning, lifting LFM2.5-350M from 22.6% to 29.7%.
- 31.9% of JSON tasks passed, up 13.9 percentage points from the 18.0% baseline.
- 29.7% of bare-list tasks passed, a 13.1-point improvement from 16.6%.
- About 500 samples and 100 GRPO steps fit the training run onto a free-tier 16 GB GPU.
- Only 6 million parameters, or 1.66% of LFM2.5-350M, were trained through the LoRA adapter.
- 29.7% brings the 350M model within 3.45 points of Qwen3.5-2B’s 33.15% score.
Risks & concerns
- 7,331 required-field failures remained after tuning, up from 7,228 for the base model.
- 890 wrong-item-count errors followed training, compared with 738 before fine-tuning.
- 555 type mismatches remained, increasing from the baseline’s 540.
- 27.5% YAML compliance improved only 0.3 points from 27.2%.
- 1,518 ms average latency exceeded the base model’s 1,453 ms.
- 29.7% overall compliance still means 1,406 of 2,000 IFStruct samples failed.