Groundedness
Check whether claims are supported by the supplied context or reference answer.
Turn answer quality into a set of small, testable decisions instead of one vague score.
The decision
Keep the rubric dimensions separate so your team can see why an answer passed or failed.
Check whether claims are supported by the supplied context or reference answer.
Separate a useful answer from a confident response that wandered away from the question.
Ask whether the answer covered the required points without collapsing everything into one label.
A safer loop
Start with examples where reviewers disagree, measure the failure modes, and only then choose an automation threshold.
“The response cites the policy, but misses the refund window.”
Answer + reference context
Use your own rubric
Ask in parallel
Keep evidence visible
Send the generated answer, question and reference material as one state.
Ask focused yes/no, choice or score questions for every quality dimension.
Store the verdict, confidence and review decision next to your evaluation run.
A rubric in one request
Start with examples where reviewers disagree, measure the failure modes, and only then choose an automation threshold.
{
"state": {
"question": "Can I get a refund?",
"answer": "Refunds are available within 30 days.",
"policy": "Refunds are available within 14 days."
},
"questions": {
"grounded": { "type": "noul", "instructions": "Is the answer supported by policy?" },
"complete": { "type": "score", "instructions": "How complete is the answer?", "criteria": ["missing", "partial", "complete"] }
}
}FAQ
The workflow keeps each criterion explicit and returns typed answers with probabilities. Your application can decide how to combine them and when to involve a person.
Yes. The criteria and answer options are part of the request, so the rubric can match your product or review policy.
No. Use representative examples to set thresholds, and route ambiguous or high-impact cases to review.