Topic
AppWorld Benchmark
2 stories tagged AppWorld Benchmark, newest first.
Aug 18, 2026, 6:09 PMArtificial Intelligence
IBM ALTK-Evolve Study: Agent Memory Must Match Model Capability
IBM's ALTK-Evolve study finds agent memory must match model capability, lifting gpt-oss-120b completion 16.1 points with just 5% more tokens on AppWorld.
DateSectionHeadlineSource
Aug 11
Artificial Intelligence
IBM ALTK-Evolve Beats ACE on AppWorld With Up to 85% Fewer Tokens
Hugging Face - Blog
▲