Module 4.3 · Topic 3
Building Quality Evaluation Criteria
Bottom Line Up Front: Quality is not binary. Define what "good enough" means for each work type: draft quality differs from client-facing, which differs from high-liability work. Rubrics ensure consistent evaluation and…
3.1 Defining "Good Enough" for Different Use Cases
Quality standards must match the stakes. Define thresholds explicitly:
| Use Case | Quality Threshold | Verification Intensity | Typical Output Examples |
|---|---|---|---|
| Internal draft | Conceptually sound; useful for ideation | Spot-check for logical sense | Internal memos, brainstorm docs |
| Client-facing | All facts accurate; sources verified; claims justified | Full fact-check; citation verification; peer review | Analyses, reports, proposals |
| High-liability | Verified against authoritative sources; limitations disclosed; alternatives considered | Expert review; full audit trail; signed-off by responsible party | Compliance assessments, financial advice |
| Public-facing | Error-free; no misleading omissions; defensible | Multiple reviewers; fact-check before publication | Published articles, press statements |
3.2 Rubric Design for AI Output Assessment
A rubric defines what strong, adequate, and weak output looks like. This removes guesswork and enables consistent evaluation.
- Identify evaluation dimensions: What matters for your output type? A market analysis needs accurate competitor data and sound recommendation logic. A compliance memo needs citation accuracy and precise language.
- Define performance levels: For each dimension, describe strong, adequate, and weak performance using observable criteria (e.g., "all sources verified" vs. "most verified" vs. "not checked").
- Weight dimensions by impact: Fact accuracy might be critical (40%), reasoning soundness important (35%), presentation secondary (25%). Weightings prioritize verification effort.
- Score and track: Assign scores per dimension. Over time, identify which models or prompts produce stronger output.
3.3 Quantitative and Qualitative Quality Measures
Measure quality on multiple dimensions. A factually perfect but poorly reasoned analysis is weak; an insightful argument is unreliable if facts are wrong.
- Fact accuracy rate: Percentage of verifiable claims that are correct. Track per output type.
- Citation precision: Percentage of sources that exist and support the cited claim.
- Reasoning soundness: Does the argument follow logically? Are conclusions justified?
- Completeness: Does output address the full scope? Does it flag limitations?
- Usability: Can colleagues use this without extensive rework?
3.4 Continuous Quality Monitoring
Track how AI quality evolves as you change models or refine prompts. Set up tracking using AI itself.
Logic behind this approach:
Use AI to generate a tracking template and monitor outputs quarterly. The goal is to surface trends so you can act on them.
Sample prompt:
What to expect in reply:
A ready-to-use CSV template with scoring guidance. After 3–4 months, you will have enough data to identify which models and prompts produce stronger output.
Quality starts with how you ask. For how prompt design affects output quality from the start, see Module 2.1.