01 / Evaluate

Answer Quality Lab

Turn answer quality into a set of small, testable decisions instead of one vague score.

Answer Quality Lab
3 / Answer + reference context
01Groundedness
0.96
02Relevance
0.88
03Completeness
0.62
Typed verdict grounded · incomplete

The decision

One answer can fail in three different ways

Keep the rubric dimensions separate so your team can see why an answer passed or failed.

01

Groundedness

Check whether claims are supported by the supplied context or reference answer.

02

Relevance

Separate a useful answer from a confident response that wandered away from the question.

03

Completeness

Ask whether the answer covered the required points without collapsing everything into one label.

A safer loop

Validate the judge before trusting the score

Start with examples where reviewers disagree, measure the failure modes, and only then choose an automation threshold.

Create a small evaluation set. Include clear passes, clear failures and ambiguous cases.
Compare against people. Review disagreement by rubric dimension instead of averaging it away.
Route uncertain cases. Send low-confidence or high-risk decisions to a human review queue.
Answer + reference context

“The response cites the policy, but misses the refund window.”

Answer + reference context

→
Typed verdictcase 18

Use your own rubric

Ask in parallel

Keep evidence visible

01

Bring the answer

Send the generated answer, question and reference material as one state.

02

Name the checks

Ask focused yes/no, choice or score questions for every quality dimension.

03

Connect the result

Store the verdict, confidence and review decision next to your evaluation run.

A rubric in one request

A rubric in one request

Start with examples where reviewers disagree, measure the failure modes, and only then choose an automation threshold.

A rubric in one request
POST /v1/evaluate
{
  "state": {
    "question": "Can I get a refund?",
    "answer": "Refunds are available within 30 days.",
    "policy": "Refunds are available within 14 days."
  },
  "questions": {
    "grounded": { "type": "noul", "instructions": "Is the answer supported by policy?" },
    "complete": { "type": "score", "instructions": "How complete is the answer?", "criteria": ["missing", "partial", "complete"] }
  }
}

FAQ

Questions before you evaluate

Is this the same as asking an LLM for a score?+

The workflow keeps each criterion explicit and returns typed answers with probabilities. Your application can decide how to combine them and when to involve a person.

Can I use my own evaluation rubric?+

Yes. The criteria and answer options are part of the request, so the rubric can match your product or review policy.

Should every result be automated?+

No. Use representative examples to set thresholds, and route ambiguous or high-impact cases to review.