Thursday, August 27, 2026
Tech Beat
Aug 14, 2026, 2:50 PMArtificial Intelligence

Kog Targets 10x Faster LLM Inference on AMD and Nvidia GPUs

French startup Kog targets 10x faster large-model inference by September after its 3,000 TPS GPU demo generated 200 business leads on existing hardware.

A silicon chip in a press releases a torrent of light, symbolizing Kog extracting more AI inference from standard GPUs.
Listen to this briefingAudio briefing

Summary

French startup Kog said on August 14, 2026 that software can extract faster AI inference from existing datacenter GPUs, contrasting with Cerebras, whose purpose-built chips received a warm May IPO debut. Kog’s May Hacker News preview ran its open-sourced, purpose-built Laneformer 2B model, about 2 billion parameters, at 3,000 single-request tokens per second on AMD MI300X and Nvidia H200 GPUs, but excluded laptop GPUs. CEO and solo founder Gaël Delalleau said the demo generated 200 tangible business leads. Because prospects would not fine-tune small models, Kog shifted to larger models while retaining its 30x faster LLM inference ambition. Targets include professional software engineering users facing hours-long Claude Code waits, Anthropic’s premium-priced Fast Mode market, and design partners whose prompt-generated games and apps could earn more through the Kog Inference Engine.

Unlike French rival ZML, whose hardware-agnostic software bypasses Nvidia CUDA across competing chips, Kog likens its low-level GPU acceleration to Stanford University’s Hazy Research. Delalleau, an École Polytechnique solid-state physics graduate, offensive cybersecurity veteran and four-time DEFCON CTF finalist, applies reverse engineering down to assembly and binary code, arguing newer GPUs’ rising memory bandwidth can improve decoding. His previous startup was TechCrunch50 2009 alum Stribe, and former co-founder Kamel Zeroual’s Varsity VC co-led Kog’s seed round. The hands-on method requires Kog’s 11-person team to spend weeks or months on each new GPU, limiting coverage. Kog eventually wants agent-based pipelines supporting more chips and models, aided by Scaleway, Bpifrance, French Tech 2030 and European sovereignty efforts. Delalleau expects its first major model at 10x speed in September, enabling traction demonstrations and a Series A raise, but LLM-scale validation remains essential.

Positives

  • Laneformer 2B reached 3,000 single-request tokens per second on standard AMD MI300X and Nvidia H200 datacenter GPUs.
  • Kog’s May Hacker News preview produced 200 tangible business leads for its software-based acceleration approach.
  • Laneformer 2B is open sourced, making Kog’s roughly 2 billion-parameter demonstration available for outside examination.
  • Scaleway, Bpifrance and French Tech 2030 support Kog as Europe pursues greater control over AI models and infrastructure.
  • September is Kog’s target for accelerating its first major model by 10x before demonstrating traction and pursuing Series A funding.

Risks & concerns

  • Kog’s 30x LLM inference claim has only been demonstrated with Laneformer 2B, not a major large language model.
  • Kog’s preview did not extend its acceleration benefits to laptop GPUs.
  • Prospective customers were unwilling to fine-tune small models, forcing Kog to redirect development toward larger models.
  • Kog’s 11-person team may spend weeks or months optimizing each new GPU, sharply limiting near-term hardware coverage.
  • Kog must validate 10x acceleration on a major model before demonstrating customer traction and seeking Series A funding.
Primary sourceTechCrunchhttps://techcrunch.com/2026/08/14/kog-is-going-deeper-to-squeeze-more-inference-out-of-gpus/
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceAug 27

OpenAI Brings ChatGPT Ads to India With 50 Brands, ₹725 Daily Floor

Artificial IntelligenceAug 27

Nvidia Nears $12.9 Billion Hugging Face Acquisition Amid Conflicting Reports

Artificial IntelligenceAug 27

OpenAI Expands Brazil Presence to Support Nationwide AI Adoption