ended9월 5일· 1 sources
800 Million Tokens for $36: Benchmarking DSH + DeepSeek V4 Pro on a Real Codebase
Why it matters
There is a problem with most benchmarks for coding agents: they are not software development. They are useful, of course. Give an agent an issue, run a test suite, check whether the patch passes. SWE-...
1
Sources
+0
24h
—
Growth
15d
Active