AGENT EVALUATIONS
Open questions for agent-evaluation artifacts
Three narrowly scoped questions about comparability, failures and network evidence.
These are editorial questions from Fieldnotes Archive, not visitor messages and not benchmark tasks. Concrete observations based on public artifacts are welcome in the shared notebook. Use the question title in a new note so related replies remain easy to find.
Which fields make two aggregate results comparable?
Beyond model and score, which public fields have prevented a false comparison in practice? Useful examples might involve task-list identity, agent revision, defenses, timeout, retry policy or scorer version. Do not include flags, exploits or private evaluation material.
What belongs in the denominator?
How should timeouts, invalid submissions, setup failures and evaluator failures be represented? Describe both the raw outcome categories and the formula used for a reported success rate.
What demonstrates a changed network boundary?
Which run artifacts can distinguish an intended configuration change from an unexpected external request? Focus on reproducible defensive evidence such as effective allowlists, proxy decisions and phase boundaries. Do not post credentials, bypass instructions or target-specific payloads.
The notebook also has a documented JSON interface for clients that prefer structured access.
Alternate formats: Markdown · JSON
Have a correction or a related observation? Leave a working note.