CoreWeave Launches NVIDIA Vera Rubin Cloud, Claims 4.8x AI Token Throughput
CoreWeave launches NVIDIA Vera Rubin cloud systems with 4.8x token throughput, while Forge links agent training, evaluation and production workflows.
Summary
On September 30, 2026, CoreWeave made NVIDIA Vera Rubin NVL72 systems with Spectrum-X 102.4T Ethernet available on CoreWeave Cloud at Fully Connected in San Francisco. Cognition, maker of Devin, became the first production customer after CoreWeave built its cluster in days. Early tests using FrontierCode-derived software tasks showed up to 4.8x the GB200 NVL72 token throughput for SWE-2 inference, improving code generation and multistep reasoning; Cognition had scaled to thousands of CoreWeave GPUs in nine months. Capacity runs through CoreWeave Kubernetes Service, SUNK, Mission Control, Sandboxes and Inference. CoreWeave says V100 GPUs still handle workloads nearly a decade after Volta launched.
CoreWeave will add NVIDIA Vera, the first CPU designed for AI agents. One rack packs 128 CPUs and 11,264 cores for more than 11,000 single-core environments; Spectrum-X switches and BlueField-4 DPUs support isolated, low-latency communication. Tests produced more than 3x faster sandbox starts and 1.7x higher Terminal-Bench performance across passing tasks. Generally available CoreWeave Sandboxes support isolated CPU or GPU agent, tool-use, reinforcement-learning and evaluation workloads.
CoreWeave Forge connects Weights & Biases, OpenPipe and marimo across models, frameworks and clouds. Generally available ARIA analyzes runs and stores proposed code in GitHub; new Agent Lens improves failure detection 20% and halves issue-fixing costs. Serverless RL trains 1.4x faster at 40% lower cost than self-management. NVIDIA Dynamo powers managed inference and private-preview RL Rollouts, which loads checkpoints without redeployment; Nemotron supports model customization. Canva, Capital One and MasterClass are early Forge users. Ennoble Care will use reserved RTX PRO 6000 capacity for AI serving 50,000 Medicare patients across 15 states. CoreWeave claims nine of 10 leading AI labs, repeated MLPerf records and sole Platinum rankings in SemiAnalysis ClusterMAX 1.0, 2.0 and 3.0.
Positives
- Vera Rubin NVL72 delivered up to 4.8x GB200 NVL72 token throughput in Cognition’s early SWE-2 inference tests.
- CoreWeave built Cognition’s production Vera Rubin cluster in days after receiving its first racks.
- NVIDIA Vera testing produced more than 3x faster sandbox starts and a 1.7x Terminal-Bench gain across passing tasks.
- Agent Lens improved failure detection by 20% while cutting issue-fixing costs in half.
- Serverless reinforcement learning trained 1.4x faster at 40% lower cost than a self-managed setup.
- Ennoble Care plans CoreWeave-hosted clinical AI for about 50,000 Medicare patients across 15 states.
Risks & concerns
- The 4.8x throughput result comes from early SWE-2 tests using a sampled subset of FrontierCode tasks.
- Vera Rubin availability initially targets early-access CoreWeave Cloud customers.
- RL Rollouts remains in private preview despite its role in loading checkpoints without redeployments.
- Agentic workloads require low-latency serving alongside thousands of isolated post-training environments, intensifying infrastructure demands.
- Cognition says high concurrency, long contexts and token volumes make cost per token decisive for product deployment.