LawQi

Module 4.1 · Topic 4

Verification and Quality Control

Bottom Line Up Front: High-risk work needs systematic verification. Low-risk work can be faster. Proportional verification means designing checks scaled to what's at stake—a marketing email needs less scrutiny than a…

4.1 Multi-Layer Verification Strategies

High-risk work needs systematic verification. Proportional verification means designing checks scaled to what's at stake.

  1. Define risk level: For this task, what's the cost of error?
    • Low risk: Marketing copy, brainstorming, draft emails (fix easily if wrong)
    • Medium risk: Financial analysis, technical documentation, customer-facing content (errors create rework or minor harm)
    • High risk: Legal documents, medical guidance, safety-critical decisions, regulatory filings (errors create liability, legal exposure)
  2. Quick-check (all tasks):
    • Read the output once. Does it feel right? Are obvious facts correct (names, dates, numbers)?
    • Spot obvious hallucinations: made-up citations, false claims presented as fact?
    • Does the tone match the context?
    • Time investment: 2–5 minutes.
  3. Detailed verification (medium-risk and above):
    • For specific claims (statistics, citations, regulatory details), verify against authoritative sources.
    • For analytical output (market analysis, financial projections), check the logic and underlying assumptions.
    • For generated text (contracts, policies), read every clause and test edge cases.
    • Time investment: 15–45 minutes depending on output length.
  4. Expert review (high-risk):
    • Have a subject-matter expert review for domain-specific accuracy.
    • Test edge cases (what breaks this contract? What happens if a customer exploits this loophole?).
    • Explicitly ask: "What am I missing that only experience would catch?"
    • Time investment: 30+ minutes.
  5. Staged rollout (production use):
    • Don't release to all customers/users at once.
    • Pilot with a small group. Monitor for unexpected failure modes.
    • Gradually expand as confidence builds.

4.2 Red Flag Patterns and Error Recognition

Red flags are patterns that signal errors. Building a personal library of red flags lets you spot problems faster than checking every detail.

Confident vagueness
"The process typically involves several steps" is a classic filler phrase. If AI can't be specific, it's probably hallucinating or generalizing beyond what it should.
Shifted terminology mid-answer
A response starts with "vendor" and later switches to "supplier" without explanation. Suggests conceptual confusion or mixing of different contexts.
Orphaned citations
"As stated in the Smith v. Jones case…" but no case details, location, or year. Likely fabricated or misremembered from training data.
Impossible specificity
"The exact percentage of U.S. companies that use this tool is 47.3%." Where's the source? Unlikely to be exact. Red flag for made-up statistics.
Logical non-sequiturs
"Because the market grew, we should fire the team." The conclusion doesn't follow the premise. Suggests flawed reasoning.
Outdated default knowledge
References to someone's role or company status that changed in the last 6 months (e.g., "The CEO of Company X is Y" when Y left 3 months ago).

4.3 Building Quality Rubrics and Evaluation Criteria

Quality standards don't just happen. You need to define what "good" looks like and measure against it systematically.

  1. Define the task precisely: "Generate customer service email" is vague. "Generate customer service email declining a refund while preserving the relationship" is precise.
  2. Write a rubric: Create a checklist for this task.

    Example for refund decline email:

    • Tone is professional and sympathetic (not cold)
    • Explains why we decline (specific policy reference)
    • Offers alternative solution
    • No grammar/spelling errors
    • Under 150 words
  3. Test on samples: Generate 3–5 outputs. Apply the rubric. What percentage passes? If less than 80%, either adjust the prompt or lower expectations.
  4. Build feedback loops: Track which outputs fail the rubric. Ask AI to iterate on the failed outputs with specific feedback ("Too wordy; rewrite in 100 words").
  5. Document standards: Write a 1-page guide saying "For refund decline emails, we expect X, Y, Z." Share with your team so everyone knows the standard.
  6. Audit periodically: Every month, pull random generated outputs and verify they still meet the standard. AI model updates or prompt drift can cause performance to slip.

4.4 Iterative Refinement Workflows

Rather than scrapping and restarting, iterative refinement with specific feedback trains AI toward your standards faster.

Situation: AI is generating product descriptions that are too technical for your marketing site.

First prompt: "Write a product description."
Result: Too technical, uses jargon.
Iterative refinement approach:
Second prompt: "That description was too technical. Rewrite it using everyday language. Explain the benefit (what does the customer gain?) not the feature (what does the product do?)"
Result: Better, but still needs work.

Third prompt: "Better. But rewrite again: remove any term a high school student wouldn't understand. Focus on the single biggest pain this product solves for a busy professional."
Result: Clearer.

Why it works: Each iteration gives the AI specific, actionable feedback. You're not vague ("make it better"); you're specific ("remove jargon," "focus on pain not features"). This teaches faster than starting over.