AI Startups Race to Replace Transformers in Next Generation LLMs
Startups challenge transformers with sparse attention, liquid networks, diffusion text and state spaces, promising faster, cheaper and smarter AI models.
Summary
Nine years after Google’s 2017 “Attention Is All You Need” paper, transformers power every major LLM, but dense attention compares each token with every other: a 10,000-word document can require 50 million multiplications. This drives energy and context limits as reasoning models add chain of thought and agents ingest documents, code bases and model outputs. OpenAI president Greg Brockman says 2026 compute spending will reach $50 billion, while the International Energy Agency expects data center electricity use to double by 2030.
Miami startup Subquadratic, led by Justin Dangel, says SubQ’s adaptive sparse attention rivals leading search and coding models, though skeptics remain; thousands await its release. San Francisco’s Manifest AI and CTO Carles Gelada say power retention uses rolling summaries and minimal retraining; PowerCoder converts StarCoder, while Brumby is claimed to rival some Alibaba Qwen versions, targeting hours-long video and weeks-long agents. Cambridge MIT spinout Liquid AI, led by Ramin Hasani, builds LFMs from 20% transformers and 80% worm-brain-inspired liquid networks, with designer AI choosing architectures. They match Qwen and Google Gemma versions four times larger, run in Mercedes vehicles or on a $50 Raspberry Pi, are free to organizations below $10 million revenue, and have 34 million downloads.
Palo Alto’s Inception, founded by Stanford researcher Stefano Ermon after 2024 work with two colleagues, uses diffusion to generate many tokens simultaneously. Its first model matched OpenAI’s 2019 GPT-2 at 10 times the speed; Inception says Mercury 2 matches some 2023 GPT-4 models, also 10 times faster. Google is testing Diffusion Gemma. Palo Alto’s Pathway, led by Zuzanna Stamirowska, replaces attention with state spaces so Dragon Hatchling can reason beyond word sequences. It solved more than 97% of over 250,000 difficult sudoku puzzles while several leading lab models solved none, suggesting nonlinguistic representations can improve efficiency and tackle reasoning transformers miss.
Positives
- Dragon Hatchling solved more than 97% of over 250,000 hard sudoku puzzles, while several leading LLMs solved none.
- Inception says Mercury 2 matches some 2023 GPT-4 models while running 10 times faster.
- Liquid AI’s LFMs match Qwen and Google Gemma versions four times larger using 20% transformers and 80% liquid networks.
- Liquid AI models run on a $50 Raspberry Pi, serve Mercedes vehicles and have reached 34 million downloads.
- Manifest AI says PowerCoder converts StarCoder to power retention with minimal retraining, while Brumby rivals some Alibaba Qwen versions.
- Subquadratic says SubQ’s adaptive sparse attention rivals leading search and coding models, with thousands awaiting release.
Risks & concerns
- A 10,000-word document can force dense attention to perform 50 million multiplications, increasing cost and electricity demand.
- OpenAI expects to spend $50 billion on computing in 2026, according to president Greg Brockman.
- The International Energy Agency predicts data center electricity consumption will double by 2030.
- Industry skepticism persists around Subquadratic’s claim that SubQ can rival mainstream LLMs using sparse attention.
- Transformers struggle with expanding context windows, chain-of-thought notes and the large information flows required by AI agents.