ended9월 5일· 1 sources

800 Million Tokens for $36: Benchmarking DSH + DeepSeek V4 Pro on a Real Codebase

Why it matters

There is a problem with most benchmarks for coding agents: they are not software development. They are useful, of course. Give an agent an issue, run a test suite, check whether the patch passes. SWE-...

1
Sources
+0
24h
Growth
15d
Active

Sources

Related Issues