ended7월 29일· 1 sources
Handbook.md shows that long policy documents do not reliably govern agents
Why it matters
Computer Science > Artificial Intelligence [Submitted on 28 Jul 2026] Title:HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following View PDF HTML (experimental)Abstract:Language-model ...
1
Sources
+0
24h
—
Growth
53d
Active