ended9월 9일· 1 sources

The state axis: why agent benchmarks keep measuring amnesiac models

Why it matters

I keep hammering the point that any coding-agent score is model + harness, not model alone. Same context-carryover rules, same note convention, same tool loop, same judge, or the comparison is garbage...

1
Sources
+0
24h
Growth
3d
Active

Sources

Related Issues