DeepSeek Harness Challenges Claude Code as V4 Pro API Prices Jump
DeepSeek launches MIT licensed Harness and V4 Pro for agents, adds OpenAI interfaces, and raises API prices from Aug. 16 with peak and off-peak rates.
Summary
On Aug. 13, 2026, DeepSeek released DeepSeek-V4-Pro-0813 across web Expert Mode, mobile and API, plus MIT-licensed DeepSeek Harness v0.1, or dsh, in developer preview. Built on Cordis, Harness makes models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration and interfaces replaceable plugins. It supports DeepSeek, Anthropic, OpenAI and compatible endpoints, while editing repositories, running shells, searching, planning, delegating and enforcing approvals. Unlike established Claude Code and Codex, it has no documented hosted agents or finished GitHub PR workflow, and warns of breaking changes. GitHub showed 27,500 stars and 2,000 forks.
V4-Pro, previewed April 24, has 1.6-trillion parameters with 49 billion active per token; V4-Flash has 284 billion and 13 billion active, both supporting one-million-token contexts. The unchanged deepseek-v4-pro identifier adds enhanced agent performance, native OpenAI Responses API, Codex integration, tool calling, JSON and Anthropic-format support. Non-think, Think High and Think Max tune reasoning. Company-reported V4-Pro-0813 scores are 87.9, 74.1, 71.1 and 67.2 on Terminal Bench 2.1, Toolathlon-Verified, DSBench-FullStack and DSBench-Hard; Fable 5 scores 77.9 and 77.2 on the middle two. Public Code Agent tests used Harness minimal mode.
At 16:00 UTC Aug. 16, flat API pricing becomes peak pricing at 01:00 to 04:00 and 06:00 to 10:00 UTC, with other hours half the new peak rate. Per million cache-miss input and output tokens, Flash rises from $0.14 and $0.28 to $0.22 and $0.66 off-peak or $0.44 and $1.32 peak; Pro rises from $0.435 and $0.87 to $0.66 and $1.98 or $1.32 and $3.96. Cache hits climb from $0.0028 to $0.007 or $0.014 for Flash, and $0.003625 to $0.022 or $0.044 for Pro. Combined uncached input and output costs reach $0.88 or $1.76 for Flash and $2.64 or $5.28 for Pro. Reuters estimates increases from 50% to above 1,100%.
Positives
- MIT-licensed Harness lets developers replace nearly every agent component, including models, tools, sessions, sandboxes, filesystems and orchestration.
- V4-Pro and V4-Flash support one-million-token contexts with 49 billion and 13 billion parameters activated per token, respectively.
- Native OpenAI Responses API and Anthropic-format support reduce integration work, while Codex receives one-click V4-Pro setup.
- V4-Pro-0813 scored a company-reported 87.9 on Terminal Bench 2.1 and 74.1 on Toolathlon-Verified.
- Hybrid attention reportedly uses 27% of V3.2’s single-token inference FLOPs and 10% of its KV cache at one million tokens.
- Harness attracted roughly 27,500 GitHub stars and 2,000 forks by Aug. 13, although DeepSeek cautioned these were snapshot figures.
Risks & concerns
- Harness remains a developer preview that explicitly warns of compatibility-breaking changes, limiting its readiness for enterprise production deployments.
- DeepSeek Harness lacks documented managed background agents and finished GitHub pull-request workflows already available through Claude Code and OpenAI Codex.
- V4-Pro’s combined uncached input and output cost rises from $1.305 to $2.64 off-peak and $5.28 during peak hours.
- Cached V4-Pro input jumps from $0.003625 per million tokens to $0.022 off-peak and $0.044 peak.
- Reuters calculated increases ranging from 50% to more than 1,100%, depending on model, token category and usage time.
- Public Code Agent results used Harness in minimal mode, so they measure the model and execution environment rather than model performance alone.