Skip to content

DeepSeek v4 pro 0813 released

3 pointsalexwwangdiscuss
On HN

Its benchmarks are comparable to Fable 5, with each having its strengths and weaknesses, resulting in a near tie.

Terminal Bench 2.1: 72.1 → 87.9, surpassing Opus 4.8's 85.0, and only 0.1 away from Fable 5's 88.0.

Cybergym: 52.7 → 83.3, surpassing Opus 4.8's 78.3, and slightly higher than Fable 5's 83.1.

DeepSWE: 12.8 → 62.7, surpassing Opus 4.8's 58.0.

AutomationBench: 12.8 → 31.8, surpassing Opus 4.8's 27.2 and Fable 5's 29.1.

Agents' Last Exam: 16.5 → 25.7, catching up with Opus 4.8.

The change in DeepSWE is particularly striking; the Preview score was only 12.8, but by August 13th it had reached 62.7.

The official version of DS4P features significant improvements in agent capabilities, and harness functionality is currently in testing and is said to release soon.

The price remains unchanged, but the official documentation page is no longer accessible.

Comments

No comments yet.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.