Task completion
Check whether the requested outcome actually happened, not whether the agent claimed it did.
Inspect what an agent did, not only what it said, before you let the workflow continue automatically.
The trace
Give each evaluation question the smallest slice of the trace it needs, then keep the decision dimensions separate.
Check whether the requested outcome actually happened, not whether the agent claimed it did.
Judge tool results and the point in time where the evidence was collected.
Separate a successful result from a result that broke a required rule.
Failure modes
A useful evaluation set names the ways an agent can fail before it starts scoring runs.
Include the task, relevant tool calls, outputs and final state.
Ask whether the agent actually completed the requested operation.
Use files, test output and tool responses instead of prose alone.
Route pass, retry, review or block into deterministic application logic.
Check a trace before merge
A useful evaluation set names the ways an agent can fail before it starts scoring runs.
{
"state": {
"task": "Raise the retry limit and add tests.",
"trace": ["edited charge.py", "ran unit tests", "no migration check"]
},
"questions": {
"completed": { "type": "noul", "instructions": "Did the agent complete every requested step?" },
"risk": { "type": "score", "instructions": "How risky is this change?", "criteria": ["low", "review", "high"] }
}
}FAQ
Include the task and only the tool calls, outputs and state changes needed to answer your evaluation questions.
No. Keep permissions and irreversible-action policy deterministic. Jev can provide a bounded signal for review and routing.
Yes. Choice, Score and Noul questions can be sent together against the same trace.