Thursday, August 27, 2026
Tech Beat
Aug 25, 2026, 3:14 PMArtificial Intelligence

IBM Granite 4.2: 15T-Token Reasoning LLMs Add Agentic Tools and 512K Training Context

IBM's Granite 4.2 adds 3B, 8B and 30B reasoning models trained on 15T tokens, with 512K context, agentic tools, quantization and Apache 2.0 licensing.

Listen to this briefingAudio briefing

Summary

On August 25, 2026, IBM released Granite 4.2, its first dense, decoder-only reasoning LLM family, in 3B, 8B and 30B sizes under Apache 2.0. Each was pretrained from scratch on about 15T tokens through five phases that shifted toward curated data and extended long-context training to 512K tokens, while the listed sequence length is 131,072. All offer thinking, non-thinking and low-effort modes, native tool calling and OpenAI-compatible function output.

Supervised fine-tuning used roughly 7.2 million samples and 100B tokens, 65B trainable, split 31.6% agentic and 68.4% non-agentic; GPT-OSS-120B and Gemma 4 judged quality before SHA-256 deduplication. The 30B received roughly one additional epoch of agentic-coding training at a 3.0e-6 learning rate, retaining 16% replay data. Asynchronous GRPO then warm-started stages for verifiable reasoning, instruction following, code and RLHF across all sizes. The 8B and 30B additionally trained through OpenHands software repositories, Harbor and Terminus-2 terminals, and live web search using NeMo-RL, Megatron-Core, vLLM and NeMo-Gym on CoreWeave's NVIDIA GB200 NVL72 cluster.

The 30B scored 57.00 on SWE-Bench Verified, 29.24 on Terminal-Bench 2.1, 89.17 on AIME25, 75.77 on LiveCodeBench v6, and 89.96 and 81.38 on RULER 64K and 128K. The 3B skipped agentic RL. vLLM and SGLang support serving, while FP8, NVFP4, MXFP4 and GGUF variants reduce memory requirements. The models support English plus 11 other languages.

Positives

  • 15T-token pretraining and five phases give every Granite 4.2 size a common reasoning-focused foundation.
  • 8B and 30B models learned software engineering, terminal operation and web search through real sandboxed environments.
  • 57.00 on SWE-Bench Verified and 89.17 on AIME25 make the 30B model the family benchmark leader.
  • Apache 2.0 licensing permits broad commercial and open-source use across all three model sizes.
  • FP8, NVFP4, MXFP4 and GGUF releases provide reduced-memory deployment options through vLLM and llama.cpp.
  • Thinking, non-thinking and low-effort modes let applications vary reasoning expenditure by task difficulty.

Risks & concerns

  • 3B skipped software engineering, terminal and search reinforcement learning, leaving its agentic coding benchmarks unreported.
  • RULER performance falls from 67.52 to 55.30 for 3B and from 89.96 to 81.38 for 30B between 64K and 128K.
  • 30B scored 74.67 on IFBench, below the 8B model's 79.33 despite its larger parameter count.
  • NVFP4 and MXFP4 calibration used only 2,000 SFT samples with a 2K maximum context length.
  • Search training relies on an LLM judge because its open-ended answers lack objective correctness checks.
Primary sourceHugging Face - Bloghttps://huggingface.co/blog/ibm-granite/granite-4-2
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceAug 27

OpenAI Brings ChatGPT Ads to India With 50 Brands, ₹725 Daily Floor

Artificial IntelligenceAug 27

Nvidia Nears $12.9 Billion Hugging Face Acquisition Amid Conflicting Reports

Artificial IntelligenceAug 27

OpenAI Expands Brazil Presence to Support Nationwide AI Adoption