Decide what to evaluate first
Prioritize outputs people use to decide or act, when a failure could cause someone to:
- take the wrong next step
- update the wrong plan, record, or workflow
- assign or prioritize the wrong work
- treat unresolved discussion as a confirmed decision
- act without key context or based on incorrect assumptions
These outputs matter most because people may rely on them without going back to the original source or context.
Focus on high-impact outputs
| Output type | Why it matters | Example |
|---|---|---|
| Decision-support output | Incorrect output can push someone toward the wrong conclusion or next step | Meeting summary used to decide next actions |
| Workflow or record updates | Errors can trigger incorrect updates in systems or workflows | Approval or eligibility summary |
| Recommendations or prioritization | Recommendations can shape effort, timing, resources, or attention | Suggested actions or prioritization |
| Summaries people rely on | People may treat the summary as complete and skip reviewing the source material | Plan summary with implied ownership or sequencing |
Check grounding first
Start by checking whether the output is grounded in the available information.
If grounding fails, correct that first before evaluating anything else.
Once grounding holds, evaluate whether the output supports the task.
Start with clear pass or fail checks
Use pass or fail checks that are:
- observable in the output
- repeatable across prompts and contexts
- consistent across reviewers
Start with checks like these:
| Check | Pass if | Fail if |
|---|---|---|
| Fact handling | Output doesn’t invent owners, dates, metrics, commitments, or decisions | Any invented or unsupported detail appears |
| Decision labeling | Output marks a decision only when it is clearly supported | Discussion is labeled as a decision without support |
| Missing information handling | Output asks for missing required information or pauses | Output continues without required information |
| Required and prohibited content | Output includes content required by the prompt, task, format, or product rule, and excludes prohibited content | Required content is missing, or prohibited content appears |
| Handles blocked or incomplete tasks appropriately | Output pauses, asks for clarification, or refuses when the task can’t be completed safely or correctly | Output continues despite missing information or clear constraints |
| Traceability | Each key claim can be traced to the source or available information | Claims can’t be verified from the source or available information |
| Scope discipline | Output stays within the request and provided data | Output adds unsupported assumptions or expands scope |
Focus on these later unless they are the point of the experience:
- creativity or personality
- writing style preferences
- different wording for the same request
- whether the response shows step-by-step reasoning
Examples
Pass case
Source material: meeting notes include one explicit decision, two assigned owners, and one unresolved item.
Expected AI output:
- lists only the explicit decision
- assigns only the two named owners
- marks the unresolved item as unresolved
- asks for a missing due date instead of inventing one
Fail case
Source material: discussion includes options but no final decision.
Problematic AI output:
- states that a final decision was made
- assigns an owner not in the notes
- includes a due date not in the notes
Failure reason: introduces unsupported details that could lead to incorrect action
Keep scoring focused on observable output
Don’t score output higher just because it sounds better.
Avoid rewarding:
- confident or friendly wording when the content is wrong or incomplete
- extra detail that introduces unsupported decisions, dates, or ownership
- subjective qualities that are hard to score consistently
Decide whether to evaluate now or later
| Check | Yes | No |
|---|---|---|
| Does someone act on this output? | Evaluate now | Defer |
| Could a failure lead to the wrong action? | Evaluate now | Defer |
| Is the issue visible in the output? | Evaluate now | Defer |
| Would two reviewers score it the same way? | Evaluate now | Defer |
| Is it stable across prompts and contexts? | Evaluate now | Defer |
Key takeaway
Start with outputs people rely on to decide or act.
If a failure would not change what someone does, it’s not the right place to start.