02 / Agents

Agent Reliability

Inspect what an agent did, not only what it said, before you let the workflow continue automatically.

The trace
review required
Collect the statePass
Check the claimPass
Inspect the evidenceReview
Choose the handoffPending
Give each evaluation question the smallest slice of the trace it needs, then keep the decision dimensions separate.

The trace

A final answer is only one frame

Give each evaluation question the smallest slice of the trace it needs, then keep the decision dimensions separate.

01

Task completion

Check whether the requested outcome actually happened, not whether the agent claimed it did.

02

Evidence quality

Judge tool results and the point in time where the evidence was collected.

03

Instruction fit

Separate a successful result from a result that broke a required rule.

Failure modes

Make agent failures legible

A useful evaluation set names the ways an agent can fail before it starts scoring runs.

01

Collect the state

Include the task, relevant tool calls, outputs and final state.

02

Check the claim

Ask whether the agent actually completed the requested operation.

03

Inspect the evidence

Use files, test output and tool responses instead of prose alone.

04

Choose the handoff

Route pass, retry, review or block into deterministic application logic.

Check a trace before merge

Make agent failures legible

A useful evaluation set names the ways an agent can fail before it starts scoring runs.

Check a trace before merge
POST /v1/agent-check
{
  "state": {
    "task": "Raise the retry limit and add tests.",
    "trace": ["edited charge.py", "ran unit tests", "no migration check"]
  },
  "questions": {
    "completed": { "type": "noul", "instructions": "Did the agent complete every requested step?" },
    "risk": { "type": "score", "instructions": "How risky is this change?", "criteria": ["low", "review", "high"] }
  }
}

FAQ

Questions before you evaluate an agent

What should be included in the trace?+

Include the task and only the tool calls, outputs and state changes needed to answer your evaluation questions.

Can Jev replace deterministic permissions?+

No. Keep permissions and irreversible-action policy deterministic. Jev can provide a bounded signal for review and routing.

Can I evaluate multiple dimensions at once?+

Yes. Choice, Score and Noul questions can be sent together against the same trace.