Module 5.2 · Topic 4
Quality and Risk in AI-Generated Artifacts
Bottom Line: AI-generated work is useful but fallible. Establish review procedures that catch hallucinations, errors, and quality issues before artifacts reach clients or the public. The cost of review is far lower than…
4.1 Reviewing and Validating AI-Created Work
You need to establish review procedures that catch hallucinations, errors, and quality issues in AI-generated content before it reaches clients or the public.
- Input verification: Before requesting AI generation, verify that your input data is correct and complete. If you ask AI to generate a report analyzing Q1 sales, and your Q1 data is wrong, the report will confidently report wrong conclusions. Spot-check key figures: does the AI's input match your records?
- Output assessment: Generate the artifact. Read it carefully, looking for three categories of error: (a) factual claims that can be verified independently (sales numbers, dates, names, statistics), (b) logical consistency (does conclusion A follow from premise B, or does the AI contradict itself three paragraphs later), and (c) domain appropriateness (does the recommendation make sense in context, or is it generic advice that misses your specific situation).
- Spot-check critical claims: Do not verify every fact (that is infeasible). Instead, spot-check 5–10 critical claims: key statistics, quotes, case references, and recommendations. If the spot-check fails (you verify a statistic and it is wrong), re-examine the entire output for pattern.
- Stakeholder review: For high-stakes work (client deliverables, regulatory filings, public communications), have a human stakeholder review the AI-generated content before release. That person brings domain expertise and external perspective that catches errors you might miss.
- Document your review: Keep a record of who reviewed the work and what they verified. If an error surfaces later, this record shows you exercised due diligence. It also creates accountability: reviewers are more careful when they know their review will be documented.
4.2 Common Failure Modes in Generated Content
You need to spot characteristic errors in AI outputs—hallucinations, inconsistencies, outdated information—through targeted spot-checks.
- Hallucinations: AI confidently states false facts as if true. Spot-check by asking: Is this claim verifiable? Can I find this statistic in a reliable source? Does this quote actually appear in the cited document? Hallucinations often cluster around statistics, case citations, and quotes. If an output cites "a 2024 McKinsey study showing X," verify that study exists and actually claims X.
- Inconsistencies: AI contradicts itself within a single artifact. Page 3 recommends "always use the conservative approach"; page 7 recommends "aggressive optimization." Scan for contradictions by reading topic headers and jumping between sections. Inconsistencies suggest the AI stitched sections together without holistic review.
- Outdated Information: AI trained on data from February 2025 will not know about March 2026 developments. If an output claims "current market trends show X" and you know that X has shifted significantly in the past year, the information is stale. Use a knowledge cutoff date as a reality check.
- Contextual Misunderstanding: AI may misinterpret your domain or audience. A prompt to "create a competitive analysis" might produce generic industry comparison when you meant specific analysis of your three closest competitors. Spot-check by asking: Did the output address my specific context, or did it produce boilerplate?
- Logical Gaps: AI may skip reasoning steps or leap to unsupported conclusions. A recommendation that "we should migrate to Platform X to reduce costs" may lack the intermediate reasoning: cost savings from reduced licensing fees, reduced infrastructure, etc. If you cannot retrace the logic, ask the AI to justify each claim.
4.3 Iterative Refinement and Version Control
You need to correct AI-generated work through follow-up prompts or edits, track versions as artifacts evolve, and maintain clear records of the final version.
- Identify specific defects: When you find an error, pinpoint it: "In the revenue forecast, you assumed 8% quarterly growth. Our historical data shows 6% growth in Q1–Q3 and a sharp dip to 4% in December. Please revise the forecast using actual historical growth rates." Vague feedback ("make it better") leads to vague iterations.
- Request incremental refinement: Do not ask the AI to regenerate the entire artifact from scratch (you lose good sections). Instead, ask it to revise specific sections: "Keep the introduction and recommendations unchanged. Revise the analysis section to reflect the corrected growth rates."
- Track versions: Name versions sequentially (Report_Draft_1, Report_Draft_2). Maintain a log of what changed: "Draft 1 → Draft 2: Corrected revenue forecast, added competitive benchmark, refined executive summary tone." This log is valuable if a third party asks what changed between versions.
- Preserve the final version clearly: When the artifact is final, save it with a clear name and timestamp (Report_Final_2026-03-02). Include metadata: author (if you are claiming responsibility), review sign-offs (who reviewed it and when), and any caveats (e.g., "Forecast assumes consistent market conditions").
- Document AI involvement: If the artifact will be shared externally (client, regulator, public), document the extent of AI involvement: "Sections 1–3 were drafted by AI and reviewed by [name]. Sections 4–5 were written by [name] and reviewed by [name]." Transparency prevents downstream surprises.
4.4 When AI-Generated Work Needs Expert Review
You need to determine which AI-generated artifacts require expert human review (technical, legal, medical) before use and escalate appropriately.
| Matter Type | AI-Generated Risk | Review Requirement |
|---|---|---|
| Legal Document or Advice | AI may misinterpret law, miss jurisdictional nuance, or suggest liability exposure the author does not recognize. Errors can result in violation or malpractice. | Mandatory attorney review before any client use or filing. AI may draft; attorney must verify and take responsibility. |
| Medical Guidance or Diagnosis | AI lacks the clinical judgment to weigh patient-specific factors. Misdiagnosis or inappropriate treatment recommendations cause harm. | Mandatory physician review. AI may assist with research or explanation; diagnosis and treatment require a licensed medical professional. |
| Technical Architecture or Security Assessment | AI may overlook security risks, suggest inappropriate technologies, or recommend approaches that fail under real-world load. Failures are costly and visible. | Mandatory review by architect or security professional experienced in your domain. AI may generate options; expert must validate. |
| Financial or Investment Analysis | AI may misinterpret data, use outdated market assumptions, or suggest strategies inappropriate for the client's risk tolerance. Losses are traceable and litigable. | Mandatory review by qualified financial professional (CFO, financial analyst, or advisor with relevant credentials). AI may provide data synthesis; professional makes recommendation. |
| Scientific or Academic Writing | AI may conflate or misrepresent prior research, suggest experimental designs that lack rigor, or claim novelty for established findings. | Mandatory peer review by domain expert familiar with the literature. Standard academic review process should examine AI involvement. |
| Marketing or Communications | AI may over-claim, offend audience segments, or violate brand tone. Errors damage reputation but are reversible. | Recommended (not mandatory) review by marketing professional or brand stakeholder. Mistakes are correctable; urgency is lower. |
| Data Analysis or Visualization | AI may choose misleading visualizations, misinterpret data relationships, or report confidence in uncertain analyses. | Recommended review by data professional. Spot-check calculations; verify that visual encoding matches data (e.g., bar height correctly represents value). |
See Module 4.3 for verification frameworks that apply to validating AI-generated artifacts.