Friday, October 2, 2026
Tech Beat
Oct 2, 2026, 4:01 AMArtificial Intelligence

ServiceNow AutoSynthData Boosts Gemma Enterprise Agent Performance 35%

ServiceNow CoreAI's AutoSynthData generated about 2,000 tasks per domain, lifting Gemma Pass@1 by 7.2 points in Hybrid and 8.41 points in enterprise ITSM.

Listen to this briefingAudio briefing

Summary

On October 2, 2026, ServiceNow CoreAI detailed AutoSynthData, a pipeline that converts a target enterprise agent’s failures and a stronger teacher’s successes into training tasks composed of a system specification, user prompt and verifier. Diagnostic runs become sanitized capability cards describing the tested skill, tools, workflow, failure, correct end state and variable dimensions, without exposing original prompts, entities, trajectories or verifier details. Generated tasks must be feasible, realistic and difficult; verifiers must be consistent, sound and complete.

AutoSynthData first creates parallel, vetted target samples, then multiplies them into novel variants with distinct requests, states, entities, reference trajectories and verifiers; multiplied samples cannot seed further variants. A shared controller manages quality, coverage and dataset construction, while an environment adapter handles execution and deterministic verification. Preferred candidates are solved by the target in no more than one of three trials and by the stronger solver in at least two of three, then pass reference execution, positive and mutated-outcome negative checks, bounded critique and repair. Batch reviews curb repetition and redirect generation toward missing capabilities. Accepted teacher demonstrations support supervised fine-tuning, with reinforcement learning and repeated, moving-frontier rounds planned but untested.

Using the released EnterpriseOps Gym dataset from Malay et al., 2026, Hybrid testing paired Gemma-4-26B-A4B-it with Qwen3.8-27B. AutoSynthData generated 2,000 samples in about 18 hours; epoch 5 was best, improving mean Pass@1 by 7.2 percentage points, or 35% relatively, raising verifier success from 63.01% to 68.55%, and closing 59% of Gemma’s original gap to the reference model. In ITSM, Gemma paired with DeepSeek-V4.1-Flash; 1,994 samples took 66 hours and raised mean Pass@1 from 18.77% to 27.18%. ITSM ran before throughput optimizations and used a larger teacher, explaining its longer generation time; evidence remains limited to this controlled environment.

Positives

  • 2,000 Hybrid samples were generated in about 18 hours, demonstrating a path to training-scale synthetic enterprise datasets.
  • Hybrid mean Pass@1 rose 7.2 percentage points, a 35% relative improvement, while verifier success increased from 63.01% to 68.55%.
  • 59% of Gemma’s original Hybrid Pass@1 gap to the reference model was closed by the best checkpoint at epoch 5.
  • ITSM mean Pass@1 increased from 18.77% to 27.18% after fine-tuning on 1,994 synthetic samples.
  • Sanitized capability cards prevent generators from receiving original evaluation prompts, entities, trajectories or verifier details.

Risks & concerns

  • 66 hours were required to generate 1,994 ITSM samples, versus about 18 hours for 2,000 Hybrid samples.
  • EnterpriseOps Gym is the only controlled environment tested, so performance beyond its Hybrid and ITSM domains remains undemonstrated.
  • Reinforcement learning support and repeated moving-frontier training rounds are planned but have not been tested.
  • Weak or overly restrictive verifiers can reward incorrect behavior or reject valid solutions, requiring positive, negative and repair checks.
Primary sourceHugging Face - Bloghttps://huggingface.co/blog/ServiceNow-AI/autosynthdata
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceOct 1

OpenAI GPT-6 Astra Ultrafast Runs Up to 8x Faster on NVIDIA Blackwell

Artificial IntelligenceOct 1

Google Wins Dismissal of Chegg and Penske AI Search Antitrust Suits

Artificial IntelligenceOct 1

13,000 AI Writing Tells Expose Claude Opus 5.5 and OpenAI Astra