LawQi

Module 4.3 · Topic 3

Building Quality Evaluation Criteria

Bottom Line Up Front: Quality is not binary. Define what "good enough" means for each work type: draft quality differs from client-facing, which differs from high-liability work. Rubrics ensure consistent evaluation and…

3.1 Defining "Good Enough" for Different Use Cases

Quality standards must match the stakes. Define thresholds explicitly:

Use CaseQuality ThresholdVerification IntensityTypical Output Examples
Internal draftConceptually sound; useful for ideationSpot-check for logical senseInternal memos, brainstorm docs
Client-facingAll facts accurate; sources verified; claims justifiedFull fact-check; citation verification; peer reviewAnalyses, reports, proposals
High-liabilityVerified against authoritative sources; limitations disclosed; alternatives consideredExpert review; full audit trail; signed-off by responsible partyCompliance assessments, financial advice
Public-facingError-free; no misleading omissions; defensibleMultiple reviewers; fact-check before publicationPublished articles, press statements

3.2 Rubric Design for AI Output Assessment

A rubric defines what strong, adequate, and weak output looks like. This removes guesswork and enables consistent evaluation.

  1. Identify evaluation dimensions: What matters for your output type? A market analysis needs accurate competitor data and sound recommendation logic. A compliance memo needs citation accuracy and precise language.
  2. Define performance levels: For each dimension, describe strong, adequate, and weak performance using observable criteria (e.g., "all sources verified" vs. "most verified" vs. "not checked").
  3. Weight dimensions by impact: Fact accuracy might be critical (40%), reasoning soundness important (35%), presentation secondary (25%). Weightings prioritize verification effort.
  4. Score and track: Assign scores per dimension. Over time, identify which models or prompts produce stronger output.

3.3 Quantitative and Qualitative Quality Measures

Measure quality on multiple dimensions. A factually perfect but poorly reasoned analysis is weak; an insightful argument is unreliable if facts are wrong.

  • Fact accuracy rate: Percentage of verifiable claims that are correct. Track per output type.
  • Citation precision: Percentage of sources that exist and support the cited claim.
  • Reasoning soundness: Does the argument follow logically? Are conclusions justified?
  • Completeness: Does output address the full scope? Does it flag limitations?
  • Usability: Can colleagues use this without extensive rework?

3.4 Continuous Quality Monitoring

Track how AI quality evolves as you change models or refine prompts. Set up tracking using AI itself.

Logic behind this approach:

Use AI to generate a tracking template and monitor outputs quarterly. The goal is to surface trends so you can act on them.

Sample prompt:

Create a CSV template to track AI output quality for [output type]. Include columns for: Date, Model Used, Prompt Version, Fact Accuracy Score (0–100), Citation Precision (% correct), Reasoning Quality (weak/adequate/strong), Reviewer Notes, Actions Taken. Add 3 example rows.

What to expect in reply:

A ready-to-use CSV template with scoring guidance. After 3–4 months, you will have enough data to identify which models and prompts produce stronger output.

Quality starts with how you ask. For how prompt design affects output quality from the start, see Module 2.1.