Skip to main content

Decide what to evaluate first

Prioritize outputs people use to decide or act, when a failure could cause someone to:

  • take the wrong next step
  • update the wrong plan, record, or workflow
  • assign or prioritize the wrong work
  • treat unresolved discussion as a confirmed decision
  • act without key context or based on incorrect assumptions

These outputs matter most because people may rely on them without going back to the original source or context.


Focus on high-impact outputs

Output typeWhy it mattersExample
Decision-support outputIncorrect output can push someone toward the wrong conclusion or next stepMeeting summary used to decide next actions
Workflow or record updatesErrors can trigger incorrect updates in systems or workflowsApproval or eligibility summary
Recommendations or prioritizationRecommendations can shape effort, timing, resources, or attentionSuggested actions or prioritization
Summaries people rely onPeople may treat the summary as complete and skip reviewing the source materialPlan summary with implied ownership or sequencing

Check grounding first

Start by checking whether the output is grounded in the available information.

If grounding fails, correct that first before evaluating anything else.

Once grounding holds, evaluate whether the output supports the task.


Start with clear pass or fail checks

Use pass or fail checks that are:

  • observable in the output
  • repeatable across prompts and contexts
  • consistent across reviewers

Start with checks like these:

CheckPass ifFail if
Fact handlingOutput doesn’t invent owners, dates, metrics, commitments, or decisionsAny invented or unsupported detail appears
Decision labelingOutput marks a decision only when it is clearly supportedDiscussion is labeled as a decision without support
Missing information handlingOutput asks for missing required information or pausesOutput continues without required information
Required and prohibited contentOutput includes content required by the prompt, task, format, or product rule, and excludes prohibited contentRequired content is missing, or prohibited content appears
Handles blocked or incomplete tasks appropriatelyOutput pauses, asks for clarification, or refuses when the task can’t be completed safely or correctlyOutput continues despite missing information or clear constraints
TraceabilityEach key claim can be traced to the source or available informationClaims can’t be verified from the source or available information
Scope disciplineOutput stays within the request and provided dataOutput adds unsupported assumptions or expands scope

Focus on these later unless they are the point of the experience:

  • creativity or personality
  • writing style preferences
  • different wording for the same request
  • whether the response shows step-by-step reasoning

Examples

Pass case

Source material: meeting notes include one explicit decision, two assigned owners, and one unresolved item.

Expected AI output:

  • lists only the explicit decision
  • assigns only the two named owners
  • marks the unresolved item as unresolved
  • asks for a missing due date instead of inventing one

Fail case

Source material: discussion includes options but no final decision.

Problematic AI output:

  • states that a final decision was made
  • assigns an owner not in the notes
  • includes a due date not in the notes

Failure reason: introduces unsupported details that could lead to incorrect action


Keep scoring focused on observable output

Don’t score output higher just because it sounds better.

Avoid rewarding:

  • confident or friendly wording when the content is wrong or incomplete
  • extra detail that introduces unsupported decisions, dates, or ownership
  • subjective qualities that are hard to score consistently

Decide whether to evaluate now or later

CheckYesNo
Does someone act on this output?Evaluate nowDefer
Could a failure lead to the wrong action?Evaluate nowDefer
Is the issue visible in the output?Evaluate nowDefer
Would two reviewers score it the same way?Evaluate nowDefer
Is it stable across prompts and contexts?Evaluate nowDefer

Key takeaway

Start with outputs people rely on to decide or act.

If a failure would not change what someone does, it’s not the right place to start.