ended9월 9일· 1 sources
The state axis: why agent benchmarks keep measuring amnesiac models
Why it matters
I keep hammering the point that any coding-agent score is model + harness, not model alone. Same context-carryover rules, same note convention, same tool loop, same judge, or the comparison is garbage...
1
Sources
+0
24h
—
Growth
3d
Active