Thursday, August 27, 2026
Tech Beat
Aug 24, 2026, 9:00 AMArtificial Intelligence

Children Outlearn AI Using 100,000 Times Less Language Data

Children learn grammar from 10 million to 30 million words, while AI consumes trillions, spurring a search for smaller, more efficient language models.

Listen to this briefingAudio briefing

Summary

Four years after ChatGPT, Claude, DeepSeek and OpenAI’s GPT models can converse fluently, but may process 100,000 times more words than humans need to acquire a first language. Toddlers usually form grammatical sentences after 10 million to 30 million words; a preteen in a language rich home may hear 100 million, rising with literacy to roughly 300 million by age 20. Meta’s Llama 3.1, released in 2024, pretrained on 15 trillion tokens; Ethan Gotlieb Wilcox estimates frontier models may use 10 times more. Easily available internet training data could be exhausted as early as the 2030s.

BabyLM, created after Alex Warstadt and Leshem Choshen discussed the idea in August 2022, tests models on developmentally plausible corpora of 100 million words or 10 million for its toddler track. The 2024 winner, GPT-BERT, combined next token prediction with masked word recovery and beat Meta’s Llama 2 70B on one benchmark using about 15,000 times less pretraining data. Yet BabyLM systems remain well below commercial LLMs; many cannot generate text, and ordering inputs from simple to complex did not deliver expected gains.

Embodied learning is the next test. Michael C. Frank’s SAYCam recorded three babies for two hours weekly from six months to two and a half years; in 2024, Brenden Lake’s team used 61 hours to teach object and word associations without presumed innate biases, but not two year old ability. Uri Hasson’s newer data covers the first 1,000 days of 17 children for 12 hours daily. Researchers are testing whether vision, hearing, active exploration and social learning can close the gap; Meta has joined a baby headcam challenge, while Q Labs launched NanoGPT Slowrun in March 2026. Success could cut training costs, widen university access, support Czech, Norwegian and Sami models, and enable controlled tests of language acquisition.

Positives

  • GPT-BERT beat Meta’s Llama 2 70B on one BabyLM benchmark despite using about 15,000 times less pretraining data.
  • Brenden Lake’s 2024 model learned object and word associations from 61 hours of SAYCam footage without presumed innate learning biases.
  • Uri Hasson’s data set records 17 children for 12 hours daily across their first 1,000 days.
  • Meta researchers joined a new benchmark and challenge focused on training models with baby headcam footage.
  • More efficient training could help universities and languages including Czech, Norwegian and Sami build capable models with limited data.

Risks & concerns

  • Llama 3.1 consumed 15 trillion pretraining tokens, while frontier systems may use 10 times more.
  • Easily available internet training data could run short as early as the 2030s.
  • Many BabyLM systems cannot generate text and remain substantially less capable than commercial language models.
  • Curriculum learning, which orders training material from simple to complex, failed to produce the expected BabyLM gains.
  • Models trained on children’s video still lack their active exploration, sensory continuity and socially informed learning.
Primary sourceArtificial intelligence – MIT Technology Reviewhttps://www.technologyreview.com/2026/08/24/1141740/kids-machines-language-learning/
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceAug 27

OpenAI Brings ChatGPT Ads to India With 50 Brands, ₹725 Daily Floor

Artificial IntelligenceAug 27

Nvidia Nears $12.9 Billion Hugging Face Acquisition Amid Conflicting Reports

Artificial IntelligenceAug 27

OpenAI Expands Brazil Presence to Support Nationwide AI Adoption