What Must Be Taught
The people responsible for deciding what lawyers must know about AI have a problem: they are not sure they know it themselves.
Reasonable Doubts · 18 chapters · ~2 min each
Chapter One (~2 minutes)
The Docket
The commission had planned to spend its first sitting adopting an agenda. Instead Justice Ada Klein put a stack of decisions on the table and asked that they be read aloud, because she wanted the room to feel the thing before it named it. So they went around: a New York affidavit citing cases that did not exist; a Wyoming patent brief that drew twelve thousand dollars in sanctions and a pointed note about the firm's absent verification process; a national firm sanctioned on appeal in California; the Massachusetts fabrications, the Fifth Circuit opinion setting guidance for every court beneath it.1 Grace Okonkwo read three and stopped. "These aren't exotic," she said. "It's the same act every time. Someone trusted an output and filed it."
That was the pattern Klein wanted on the record: that the failures sorted into a small number of exposures, malpractice by AI error, confidentiality breached at the point of entry, unauthorized practice, the waiver of privilege, the violated standing order, and that they compounded, one bad output metastasizing across every matter that reused it.2 But the discomfort in the room was not the liability. It was the biography. Victor Hensley, who prosecutes these people for a living, said it plainly. "Not one of them was ever taught this. Not in school, not on the bar exam, not in a CLE, not by a supervising partner. We are disciplining lawyers for failing a test we never set."
Prof. Daniel Estrada objected that the duty had existed all along, competence having hardened over just a few years from a genial suggestion of technological awareness into a settled expectation that a lawyer understand the tools of the practice.3 "It existed," Hensley said. "It was never delivered. A duty nobody teaches is a trap, not a standard."
Klein steered them toward the mechanism, and here the room's quietest member spoke. Margaret Osei used none of these tools and was, for that reason, the one who could say cleanly why they failed. The systems are fluent by construction and accurate only when the statistics happen to land on the truth; they are pattern-matching engines that produce a plausible sentence with the same confidence whether it is sound or fabricated, so the fluency that makes the output persuasive is exactly what makes it dangerous.4 "Every lawyer on that docket read something that read well," she said, "and 'read well' was the whole of their verification."
Someone asked how she knew this if she never used them. "Because I had to decide not to," she said. "That took more study than clicking accept." It was the first appearance of what became the commission's structural discovery, which was that declining well was not the absence of a competency but a competency of its own, and that it had gone untaught alongside everything else. No syllabus had ever asked a student to refuse a tool for stated reasons and document the alternative. The docket punished the users; it had never reached the abstainers, because abstention was assumed to need no teaching at all.
By the end they had a method, if not an agenda. Klein proposed the test the whole commission would live under: no competency counted unless it could be turned into a class, an assignment, and a grade. "If we can't teach it, examine it, and score it," she said, "we can't require it. And if we can't require it, we've no business disciplining anyone for lacking it." She looked at the pile. "Every one of these is a final exam we administered without ever offering the course. We are going to build the course."
Plate One
Chapter Two (~2 minutes)
Comment 8 and What It Never Meant
The document under examination was a single sentence. Raymond Tull read the comment to Model Rule 1.1 aloud the way one reads a will after the funeral: the duty to keep abreast of changes in the law and its practice, including the benefits and risks associated with relevant technology. "Some forty states adopted this," he said. "None of them defined 'relevant technology,' and none said what 'keep abreast' costs in hours. We have run a decade of enforcement off a subordinate clause."
Klein wanted it examined as drafting, not as scripture. The comment had done real work. It had moved competence from optional to mandatory and made a tool's capabilities and limitations something a lawyer owed the client rather than a hobby the enthusiast indulged.1 But everything built on top of it, the formal opinion and the state variations, described a lawyer who uses the tool. "Read literally," Okonkwo said, "the duty attaches to the user. The lawyer who advises on a model-drafted contract, or cross-examines a model-made exhibit, sits outside the sentence. That can't be right, but that is what it says."
Estrada found the deeper flaw temporal. Comment 8 has no trigger until something has already gone wrong; it is discovered, always, in the past tense, at a sanctions hearing, after the client has been hurt. "It enforces only after harm," he said. "That's not a standard of care. It's an autopsy protocol." Hensley agreed with a prosecutor's bleakness; he had never once charged the comment prospectively. It names a duty of literacy, to understand the technology well enough to protect a client who depends on it,2 and then waits for the failure that proves the literacy was absent.
Osei pressed the part the others kept sliding past: the duty is continuing, not discharged. The comment says keep abreast, present tense and indefinitely, and the tools change under the practitioner every few months, so a competence certified in spring is stale by autumn.3 "You cannot satisfy this once," she said. "Which means you cannot satisfy it with a single seminar, which is exactly what every one of you is planning to build." The room was quiet, because she was right, and because she was the member least likely to have said it in her own defense.
There was a second duty folded into the sentence that no one had been enforcing at all. The same literacy that lets a lawyer use a tool is what lets her advise a client's board on governing one: which regulatory regimes now reach the system, what the emerging duties of transparency and human oversight require, where the exposure sits when a model's decision harms someone.4 "Comment 8 obliges the lawyer to counsel on AI," Klein said, "using knowledge Comment 8 declines to specify. We have written a mandate and withheld the curriculum."
Then the test. Could "keep abreast of relevant technology" be turned into a class, an assignment, and a grade? Estrada tried and could not; the phrase was a direction of travel, not an outcome you could observe in a candidate. What survived the exercise was narrower and more useful than the comment itself: not the sentence, but a translation of it into specific, testable objectives. Name the failure modes, state the confidentiality boundary, verify an authority. Each could be examined, because each described something a person either did or did not do. "The comment tells us the duty exists," Tull said, closing his copy. "It has never once told us what it looks like when discharged. That has been our job all along, and we have been leaving it to the sanctions docket to do for us."
Plate Two
Chapter Three (~2 minutes)
The Self-Audit
It was Osei who proposed that the commission sit the diagnostic before it imposed one, and Klein who insisted the results be published in aggregate so no member could quietly decline. The instrument was modest: explain, in plain terms, what these systems are and why they fail. They took it in silence and scored it together, and the silence afterward was the useful part.
The first item asked what generative AI actually is. The strong answers placed it correctly, as a subset of machine learning that generates content by predicting what comes next, probabilistic where ordinary software is deterministic, so that the same prompt can yield two different and equally "valid" answers.1 The weak answers described a search engine that occasionally lied. The gap mattered, because the whole of the duty flowed from which picture a lawyer carried in her head.
The second item asked how the thing works, and here Estrada's students would have outscored two of the regulators. The mechanism is not retrieval; it is prediction, token by token, with language turned to numbers, run through layered mathematics, turned back into words, and a deliberate seasoning of randomness that is a feature rather than a defect.2 "If you don't hold that," Osei said, "you will treat a confident answer as a looked-up answer, and you will be wrong at the worst possible moment." The third item asked why it errs, and the answer was the same fact wearing a darker coat: a model with a gap in its training fills the gap with the likeliest-looking continuation, which is why a fabricated citation looks exactly like a real one until you check it.3 Confabulation was not malfunction. It was the machine doing precisely what it was built to do.
Klein read the aggregate scores and did not soften them. The regulators, herself and Hensley and Judge Nawaz, had scored lowest as a group; the academics had done well; and Osei, who used none of it, had very nearly topped the sheet. "Put that in the record exactly as it is," Klein said. "The people who discipline, adjudicate, and set the standard understand the technology least. If that isn't published, the whole exercise is theater."
Hensley wanted to know what the number even measured, and whether a diagnostic could distinguish real understanding from the memorized vocabulary of understanding. It was a fair objection and it seeded a later chapter, because the same doubt would haunt the bar exam. But Estrada held the line on this much: the diagnostic had exposed a genuine skill deficit, not merely a rhetorical one, and skill of this kind develops only through graded practice, built in steps and held at each step until it holds, the way any competence is actually acquired rather than merely described.4 "You cannot brief your way to this," he said. "It has to be done, repeatedly, and it hasn't been."
The last item was Osei's favorite because it was the one the regulators found most alien: how a model's behavior is governed by hidden instructions and a strict internal chain of command, so the tool is not an open field but a constrained instrument whose limits a probing user can map.5 "You can learn the shape of the cage," she said, "without ever living in it. That is the whole of my method, and it scores." The argument that closed the sitting was not about the syllabus at all. It was whether to publish the fact that the commission had failed its own first test, and Klein, overruling two members who thought it undignified, ruled that they would. "A commission that hides its own diagnostic," she said, "has no standing to grade anyone else's."
Plate Three
Chapter Four (~2 minutes)
Nobody Owns Competence
Having failed itself, the commission drew a map, because Klein suspected the profession's unreadiness was structural before it was personal. She put five boxes on the board: the state supreme court regulates; the bar examiners test; the accreditor accredits; the CLE board certifies hours; the firm trains. Then she asked each member whose job entry-level AI competence actually was, and watched every finger point at a different box.
"That is the finding," she said. "Not that anyone dropped it. That each of us assumed another had it." Tull conceded that accreditation site visits contained no technical question on the subject; Okonkwo, that the bar exam had never tested it; Hensley, that he could only reach conduct after the fact. The duty existed at every level and was owned at none, an orphan in a building full of guardians, which is how a decade of graduates had reached practice with the standard of care rising above their heads.1
Osei made the room look outward before it looked in. The clients these lawyers advise were already governed: by the EU AI Act's risk tiers arriving on their deadlines, by the NIST framework's govern-map-measure-manage discipline, by a certifiable management standard their multinationals were being audited against.2 "The world has built an entire apparatus for governing this," she said, "and our profession, whose job is governance, cannot yet teach a new lawyer to read it. We are the least-regulated actor in our own subject."
That was the reframing. The competence in question was not merely how to operate a tool; it was how to advise on one. A lawyer counseling a board needs to know which regime reaches a given system, what the duties of transparency and human oversight require, where liability sits when a model's decision harms someone.3 And inside a firm, the same knowledge governs the firm's own use, running to the written policy defining which tools are sanctioned, the verification gates on high-stakes extractions, and the audit trail that proves competence when a bar counsel later asks.4 "Two audiences," Estrada said, "one body of knowledge. The client's governance and the firm's are the same course taught twice."
The pedagogical test cut hard here, because governance is where a syllabus is most tempted to bloat. Estrada refused to let "understand global AI regulation" onto a first-year list; no one could turn that into a fair exam item, and a graduate who could recite the EU tiers had learned trivia, not judgment. What survived was narrower and testable: not the content of every regime, but the skill of locating the regime that applies and reading a governance policy against it, a competence you could set as an assignment and grade against an answer key. "Owning competence," Klein said, "starts with one of us owning this rung. The map is useless if it only proves the orphan exists."
By the close they had assigned custody, provisionally, of each rung to a member, and had discovered in doing so that the assignment forced precision the argument never had. You cannot hand someone a duty to teach without saying what the duty is. "That," Tull said, "may be the only thing this commission does that matters. Not the syllabus. The moment each institution admits which part is its own."
Plate Four
Chapter Five (~2 minutes)
What a 1L Must Know
Estrada arrived at the first rung with a grievance and a constraint. The grievance was the wish list the commission had sent him, a page of everything a lawyer should ideally understand. The constraint was fourteen weeks, shared with contracts and the rest of a first-year life. "You have given me a curriculum for a doctorate," he said, "and a seminar's worth of room. Tell me what a 1L must know, and I will tell you what I can teach, examine, and grade. They will not be the same list, and the gap is the whole lesson."
He started by cutting toward the irreducible. A first-year did not need the transformer's internals; she needed the one fact that reorganizes every intuition after it, which is that generative AI predicts the next token from statistical patterns rather than retrieving verified fact, and is therefore variable where ordinary software is exact.1 "If they leave with only that," Estrada said, "they will distrust the right things." The second class earned its place by deepening the first: how the prediction actually runs, language to numbers to language, with sampling that makes the same question yield different answers.2 Cho asked whether that was too technical for October of first year. "It's less technical than the rule against perpetuities," Estrada said, "and considerably more useful."
The third class was the one he refused to cut, because it was the one the docket was built out of: why the machine gets things wrong. A model with a gap fills it with the likeliest-looking continuation, so hallucination is not a bug to be patched but a property to be managed, most dangerous exactly where the model is most confident and least grounded.3 "That is the single sentence that would have saved every lawyer in Chapter One," Hensley said. Estrada agreed, and noted its examinable virtue: you can hand a student a fabricated citation and a real one and ask her to say which conditions produced the fake, and you can grade the answer.
Then the two classes that made the doctrine land. A student who has only read about the tool has not met it, and the interface itself teaches something no reading conveys, which is that this is not a search box but a place where work is staged, set up and revised across several turns.4 And there is a threshold of ordinary hands-on confidence, the small competence of setting a task and getting something usable back, that a lawyer has to actually cross rather than be told about.5 "First contact is a class," Estrada said. "It cannot be a footnote."
The List B beat was Cho's insistence, and Estrada built it in from day one. The first unit that teaches a student to operate the tool must, in the same breath, teach her to recognize the matter on which she should not, noticing a live client's confidence, a rare-fact question, a stakes level that outruns her competence, and stopping. "A curriculum that teaches only 'yes' has taught half a skill," Cho said. "The 1L version of refusing is simply knowing, on day one, where the edges are."
What survived was shorter than the wish list and honest about it. Estrada could teach five things well and examine four of them; the fifth, genuine judgment about when to reach for the tool, he could model and provoke but not fairly grade in a first-year. "That is not a failure of the syllabus," he told the commission. "It is the first appearance of the thing we will keep meeting: most of this competence is judgment, and judgment is exactly what a fourteen-week course can start and cannot finish."
Plate Five
Chapter Six (~2 minutes)
The Clinic Student and the First Live Client
Lydia Cho brought a witness. The clinic's client, a tenant named in a housing matter, sat at the end of the table and answered the commission's questions about what she had understood when a law student first typed her circumstances into something. She had understood nothing; no one had asked her. "That," Cho said when the client had gone, "is where the duty actually bites. Not in a lecture hall. The first time a student holds a real person's confidence in her hands."
The clinic, Cho argued, was the rung where the abstract became irreversible, and it demanded a standard stricter than the one the commission was drafting for practicing lawyers. The first competency was confidentiality at the point of entry: a student must know, before she touches a keyboard, that vendor data policies vary wildly, that some tools retain inputs for training and grant employee access while others promise none of it, and that redaction after disclosure is no cure for disclosure.1 "A partner who breaches this has insurance and a decade of judgment," Cho said. "A 2L has a real client and no cushion. The standard has to be higher precisely because the stakes are, for her, total."
The second competency was the one students found seductive and Cho found most dangerous: the tools are genuinely powerful at document work, ingesting a file and answering questions across it, and a clinic student handed that power will reach for it on exactly the material she must most protect.2 The skill was not the capability; it was the discipline of using it only within a channel that keeps the client's confidence inside the building.
The third pushed into territory the commission had not expected at this rung. Clinic tools increasingly did not merely answer but acted, drafting and filing and moving work along, and the professional-responsibility rules were unambiguous that delegation transfers no responsibility, that a student and her supervisor remain accountable for anything an agent produces and must verify it before it becomes an act with legal effect.3 "A student cannot supervise what she doesn't understand," Cho said, "so understanding the agent is prerequisite to using one, which for most of my students means not yet."
The fourth was craft with a legal edge. When a clinic student frames a research task, the instruction must carry its own guardrails: the jurisdiction named, source validation demanded, the invention of authority forbidden. A fluent falsehood in a housing case is measured in a family's eviction, not a bad grade.4 Cho could teach this, set it as an assignment, and grade it against whether the guardrails were present.
The List B beat was the chapter's spine and the reason Cho had brought the client at all. The competency that mattered most was declining on a real file: the student who looks at this matter, this tool, this data, and says no, for stated reasons, and does the work the older way, and documents why. "I can grade that," Cho said. "I give them a live-seeming matter where the right answer is refusal, and I score whether they saw it and whether they could defend the refusal to a supervisor and to the client. A student who uses the tool brilliantly on a file she should never have uploaded fails. A student who declines, cleanly, with an alternative, passes." The standard the commission drafted for the clinic came out stricter than the one for admitted lawyers, and no one in the room could argue it should be otherwise.
Plate Six
Chapter Seven (~2 minutes)
What the Bar Exam Can Test
Grace Okonkwo came to explain why most of the commission's wish list would never survive contact with a psychometrician. "An exam item is not a good question," she began. "It is a defensible one. It must be reliable, so the same candidate scores the same twice. It must be scorable, so two graders reach the same result. And it must be valid, so that it measures the thing we claim and not something merely correlated with it. Bring me your competencies and I will tell you which ones become items and which ones die at this table."
She started with the one that seemed easiest and showed why it was hard. The commission wanted candidates to demonstrate the professional literacy the standard now requires, meaning an understanding of a tool's capabilities and limitations sufficient to protect a client.1 "I can test the knowledge component," Okonkwo said. "I can ask what a hallucination is and grade it cold. What I cannot test on a timed exam is whether this person will actually verify at midnight under deadline. Knowledge is examinable. Conduct, mostly, is not." That distinction between knowledge and judgment became the chapter's refrain.
Some competencies converted cleanly. Building an evaluation rubric was one: give candidates an AI-generated memo and a defined standard, and ask them to score it for factual accuracy, citation precision, and soundness of reasoning against observable criteria.2 "That is a beautiful item," Okonkwo said, "because the answer key is defensible. Two graders agree. It measures a real skill and it resists guessing." Analysis converted too. Hand the candidate an output and ask them to flag the gap, the buried adverse fact, the leap from premise to conclusion, treating every pattern the machine offers as a hypothesis to verify rather than a finding to adopt.3 Both were scorable because both asked the candidate to do something on the page, not to be something.
Where the list died was the hands-on skill the academics most wanted certified. Estrada argued that operating the tool, the graded and developed competence of actually using one well, belonged on the exam.4 Okonkwo was sympathetic and immovable. "A performance task requires a live system, a controlled environment, and human raters. The multistate format cannot deliver it reliably at scale. If we put it on, we cannot defend the scores, and a candidate who fails will sue us and win." What survived was a proxy: the exam could test whether a candidate recognizes good and bad practice, and could not test whether she practices well. "That is a real limit," she said. "Do not let anyone pretend the proxy is the skill."
The List B beat sharpened the whole discussion, because refusal turned out to be more examinable than use. Okonkwo could write a defensible item in which the correct answer is to decline: a fact pattern where the tool, the data, and the stakes make non-use the competent choice, and the candidate must select it and state the ground. "Judgment resists testing," she said, "but a well-built decline item comes closer than anything, because there is a defensibly right answer and a defensibly wrong one." The commission left with a shorter list than it brought and a clearer one. Most of what it wanted to require was judgment; most of what it could examine was knowledge; and the honest bar exam would test the second while saying plainly, on its own face, that it was not the first.
Plate Seven
Chapter Eight (~2 minutes)
The First Ninety Days
The associate rung produced the commission's only genuine standoff, because two institutions arrived with incompatible assumptions and neither would concede. Estrada spoke for the schools: a graduate arrives with foundations, not finish. A managing partner the commission had invited spoke for the firms: we assume a new associate can be handed a matter. "You cannot both be right," Klein observed. "A graduate is either ready to do this work in her first ninety days or she is not. Tell me what 'ready' means as a competency, and we will see who is fooling themselves."
They built the list forward from the work itself. The first was foundational prompt craft, the unglamorous discipline that clarity and specificity drive output quality and that a request stating role, context, format, and the reasoning to be shown gets a usable answer where a lazy request gets mush.1 "Firms assume this is intuitive," the partner said. "It is not," Estrada answered. "It is a taught skill, and if you assume it you will be disappointed for a quarter and blame the wrong party." The second deepened it: hard work is not one question but a sequence, facts established in one turn, a framing tested in the next, the real request built on top, which is a staircase rather than a search.2
The third and fourth were the ones the firms wrongly assumed the schools had finished. Protecting quality is a standing habit of asking the system for its sources and limits and never treating a fluent answer as an authoritative one.3 Turning classroom skill into professional application means verifying every output through real research and keeping client confidences out of any tool that falls short of the firm's security standard.4 "A school can start these," Estrada said. "Only a firm, with real matters and real supervision, can finish them. That is the ninety days. The handoff is the curriculum."
Two more the associate needed before she could be trusted with a matter. Legal work demands tools connected to authoritative sources and a verification workflow that catches errors before they reach a client: spot-check the citation, confirm the holding, validate completeness, document the check.5 And model selection is itself a competence, since a fast chat model is insufficient for analysis that bears on legal judgment, choosing the wrong tier is not a preference but a potential malpractice, and the verification owed varies by the tier you chose.6 "An associate who runs a reasoning-dependent question through a speed model," the partner admitted, "and I have several, is making an error neither of us taught her to avoid."
The List B beat resolved the standoff without either side surrendering. The competency both could agree on was knowing when a new associate should refuse, and should carry a task back up the chain rather than complete it, because the tool is wrong for it, or the data cannot go into it, or the stakes exceed what she can verify. "That," Klein said, "is the readiness marker. Not that she can use these tools. That she knows the task she should hand back." What survived the standoff was an honest division of labor: the school certifies foundations and the firm certifies application, the ninety days is the seam where they meet, and a graduate is "ready" not when she can do everything but when she reliably knows the boundary of what she should attempt alone.
Plate Eight
Chapter Nine (~2 minutes)
What a Supervising Partner Must Know
The commission had climbed the ladder expecting the training need to concentrate at the bottom. The supervising-partner rung reversed the assumption, and Klein let the reversal land. "The associate at least knows she is learning," she said. "The partner supervising her has never done this work and believes he does not need to. The highest untaught competence in this profession is not at the entrance. It is in the middle, in the people signing off."
The partner's first competency was delegation, which sounds trivial and is not: an agent, like a junior, fails on vague instruction and succeeds on precise scope, explicit success criteria, and clear exit conditions, so a partner who cannot specify a task cannot safely delegate it to a person or a machine.1 "Supervisors who delegate badly to humans," the invited partner admitted, "delegate catastrophically to agents, because the agent won't push back."
The second was oversight as an active practice rather than a final glance. Supervising agent-assisted work means checkpoints rather than a single inspection at the end, watching for the looping, the scope creep, the confident factual error and the drift from format, and intervening with concrete correction before bad work cascades.2 The third was the judgment underneath the oversight: trust is not a switch but a graduated thing, calibrated to stakes, earned on low-risk work before it is extended to high, with verification scaled from a spot-check to full traceability as the consequences rise.3 "A partner who trusts an agent the way he trusts a twenty-year associate," Osei said, "has miscalibrated in the most expensive direction. And no one ever taught him the dial exists."
The fourth and fifth lifted the rung from the matter to the firm. A supervising partner is increasingly responsible for scaling a capability across a team, moving from one fluent champion who becomes a bottleneck to templated, standardized workflows with quality thresholds and escalation protocols the whole group can execute.4 And the barrier to all of it is not technical but human, since nearly two-thirds of organizations name people and culture rather than tools as the obstacle, so the partner's real task is change management: training tiered to role, superusers, visible leadership, and the willingness to be a beginner in front of juniors.5 "The competence we are describing," Klein said, "is mostly the courage to admit incompetence in a room where you outrank everyone."
The List B beat was the one the commission had not anticipated and found most important: a partner must be able to supervise a refusal, not only a use. When an associate declines a tool for stated reasons, the supervisor has to be competent enough to evaluate whether the refusal was sound, distinguishing principled non-use from mere avoidance and backing the associate who was right to say no. "A supervisor who can only assess whether the tool was used well," Osei said, "cannot protect the associate who was right not to use it. He will punish the very judgment we are trying to teach." The pedagogical problem the rung exposed was structural: the commission could design the class and the assignment, but the people who most needed the grade were partners who would never sit for one. What survived was a competency framework the middle of the profession demonstrably lacked and had no existing mechanism to be taught. The commission recorded that as its most serious gap, because a ladder is only as strong as the hands reaching down from the rung above.
Plate Nine
Chapter Ten (~2 minutes)
The Bench
Judge Patricia Nawaz brought her own standing order and asked the commission to help her enforce it, and by the end of the sitting she had conceded she could not, a concession that cost the commission a full session and bought it more credibility than any competency it had drafted. "I ordered that parties disclose AI use in filings," she said, reading her own language. "I now cannot tell you what it requires, who it binds, or how I would detect a violation. I have been enforcing a sentence I did not understand."
The competence a judge needed, the commission found, began where the docket was already arriving regardless of anyone's policy. A judge must understand what AI-generated artifacts now are, and that the question has shifted from whether an exhibit looks authentic to where it came from, because provenance rather than appearance is the only remaining test of a thing a machine can fabricate perfectly.1 Nawaz had assumed authenticity was still a property you could inspect. It was not. Multimodal models do not merely read images, audio, and video; they generate them, whether a transcript, a voice, a signature or a photograph, at a fidelity the eye cannot catch, so "it looks real" had quietly stopped being evidence of anything.2 "I have admitted exhibits this year," she said slowly, "on a standard that no longer exists."
The chapter's hard pedagogical work was the standing order itself. Nawaz's disclosure requirement, examined as drafting, could not be complied with as written, since it did not say which uses of AI triggered disclosure, distinguish drafting assistance from generated evidence, or specify what a compliant disclosure contained. The disclosure landscape she was operating in was real, a patchwork of standing orders and certification rules and statutes, but her own contribution to it named a duty without an observable act.3 "A judge who orders disclosure," Okonkwo said gently, "has the same problem we found in Comment 8. If you cannot say what compliance looks like, you cannot find non-compliance, and your order is a gesture."
The List B beat cut against the judge's instinct. Nawaz had reached for disclosure as the answer; the commission pressed her toward verification instead. Disclosure tells the court that a tool was used; it does not tell the court whether the output is sound. The competence a bench actually needs is the ability to test the artifact, demanding provenance, chain of custody and reproducibility, rather than to rest on a party's self-report of AI involvement. "Disclosure without verification," Osei said, "just moves the credulity from the exhibit to the certification. You have to be able to check, not merely to be told." And the judge had to know the case law that was already punishing the failure, the sanctions for fabricated citations that put every filing under Rule 11 suspicion.4
What survived was narrower than Nawaz's order and more useful. The competency the commission could teach and examine was not "require disclosure" but "test provenance and verify a challenged artifact against an observable standard." Nawaz rewrote her order on the record, replacing the undefined disclosure trigger with a certification of verification tied to a specific act. The session it cost was the session in which a sitting judge said, into the transcript, that she had been enforcing something she did not understand. "That," Klein said afterward, "is the most valuable hour we have spent. Every judge in the state has that order in a drawer. She is the first to admit she couldn't defend it."
Plate Ten
Chapter Eleven (~2 minutes)
The Prosecutor's Problem
Victor Hensley put the problem on the table without softening it. "I am asked to charge, prove, and settle matters involving systems my office cannot evaluate and cannot afford to have evaluated. I do not have Meridian's engineers. I have a paralegal and a subscription. When a respondent says his tool did not do what we allege, I frequently cannot say whether he is lying." The commission had reached the rung where the competence in question belonged to the enforcer, and where its absence was itself, Hensley argued, a competence failure that belonged in the framework.
His first need was the foundational one, and he lacked it as much as anyone on the diagnostic: understanding AI reliability well enough to know what kind of error he was even alleging, whether hallucination, biased interpretation, drift or omission, because the charge depends on the failure mode and he had been charging them interchangeably.1 "If I cannot name the error," he said, "I cannot prove it, and I certainly cannot settle it for the right number."
The second was tool evaluation as a discipline he could not currently perform. To assess whether a respondent's tool was fit for the work, a prosecutor must be able to interrogate its security posture, its data handling, its accuracy on the task, its very architecture. That is the vendor-assessment competence the profession expects of adopters and had never expected of the office policing them.2 The third let him triage what he could not fully investigate: not every AI use carries equal risk, so an office without resources must classify by impact, giving the systems that touch liberty or livelihood the formal scrutiny and the low-stakes uses a lighter touch.3 "That is how I survive the caseload," Hensley said. "But it is a competence, and no one taught it to me, and my triage is currently guesswork."
The fourth was the technical fact that kept defeating him in settlement. When a respondent argues his output was an anomaly, the prosecutor needs to understand how different models process inputs: a fast model generates directly while a reasoning model deliberates, and the difference goes to whether an answer is even reproducible, which is the crux of half the disputes he sees.4 "Respondents hide in reproducibility," he said. "The same prompt, a different answer, and they call it proof of nothing. I need to know when that argument is real and when it is a shield."
The List B beat was the chapter's uncomfortable heart. Hensley's office could not evaluate the tools, and that gap was not a resource footnote; it was a competence deficiency in the enforcement arm that made the entire standard hollow. "You can write the most beautiful competency framework in the country," he told the commission, "and if my office cannot assess whether it was breached, you have written aspiration. The refusal I most need to teach is my own: the discipline to decline to charge, or to settle honestly, when I genuinely cannot evaluate the system, rather than bluffing a respondent who knows I cannot." What survived was a finding the commission had not planned to make: that the framework had to specify the enforcer's competence and the enforcer's resources as explicitly as the entrant's, because a duty no one can prove is not a standard, and an office that cannot evaluate the tool cannot pretend to.
Plate Eleven
Chapter Twelve (~2 minutes)
Teaching the Refusal
The commission had assigned Margaret Osei the verification unit expecting a short, principled lecture from the member who used none of these tools. What she delivered was the most technically demanding session of the entire body of work, and it settled the book's central claim: that competent non-use requires more understanding than casual use, not less. "You gave the abstainer the hardest material," she said, opening, "and you did it by accident. Let me explain why it was correct."
Verification, she argued, is the whole of the competence the docket was missing, and it is proportional by nature, with checks scaled to stakes, indifferent to whether the error came from a language model or a tired associate at midnight, a quick read for the trivial and expert review with a full audit trail for the consequential.1 "Everyone wants to teach 'use the tool well,'" she said. "The load-bearing skill is 'confirm the output, or refuse it.' That skill does not require you to trust the machine. It requires you to understand it well enough to distrust it precisely."
She built the unit from frameworks, not attitudes. Verification is systematic confirmation, not spot-checking: you decide in advance what you will check and how, retrieving the source directly rather than trusting a paraphrase, confirming the passage supports the claim, checking currency, looking for contrary authority, documenting it.2 In legal work the standard is the highest there is, since every citation must be checked for existence, holding, currency, and relevance, because submitting a fabricated cite is sanctionable and the duty to verify before filing is not optional.3 "This is not the abstainer's eccentricity," Klein noted. "This is the exact skill absent from every matter on our docket."
The deepest part of the session was psychological, and Osei taught it more convincingly than any user could have. Trust must be calibrated continuously and made a habit rather than a heroic act. Treat a confident conclusion as a warning rather than a comfort, default to verification, and build small routines that run on autopilot: spot-check immediately, ask what the most damaging error would be, and check for that.4 "The user's temptation is to relax as fluency rises," she said. "Mine was never to relax at all. That is why I understand the failure mode better. I have spent years not being seduced by it." And underneath it all sat the mechanism from the first diagnostic: the machine fills gaps with the likeliest-looking continuation, so it is most dangerous exactly where it is most confident and least grounded.5 "You cannot verify what you do not understand," she said. "To teach refusal I have to teach the machine more thoroughly than the person who leans on it, because I am refusing for reasons, and reasons have to be technically correct or they are just fear."
The List B claim became the commission's structural finding here. A curriculum that teaches only competent use has taught half the skill; the other half is competent declining, with reasons and a documented alternative, and it is not the easy half. It is harder, because it demands the understanding to know exactly why this tool, on this matter, should not be trusted, and the discipline to do the work another way. "The abstainer is not the person who understands least," Klein said, watching the room absorb it. "On this evidence she is the person who understands most, and we very nearly built a framework that had no place for her." What survived the session was the recognition that verification and refusal were not a footnote to the syllabus but its spine, the one competence that, taught well, made every other rung safe, and the one the profession had most completely failed to teach.
Plate Twelve
Chapter Thirteen (~2 minutes)
Who Is Qualified to Teach This?
Raymond Tull raised the question that unraveled the commission's confidence: the framework assumed instructors, and there were almost none. "We have designed a ladder," he said, "and we have not asked who is competent to stand at the top of it and teach. When I count qualified faculty in this state I run out before I fill one law school. So the vendors have volunteered, and that is its own problem."
The instructor, Tull argued, needed a genuine command of the landscape, not a product demo. A teacher must understand foundation models and providers, and know that every tool sits on a base model with a knowledge cutoff, a context window, a reasoning depth, and a provider whose interests are in the room.1 A teacher must be able to evaluate a tool stack as an ethics decision, weighing security, data flow, compliance, vendor lock-in and portability, and to teach that evaluation rather than a brand.2 "A vendor cannot teach tool-neutral evaluation," Osei said, "because the honest conclusion sometimes is 'not this one, and possibly none.' That sentence is not in their material."
The competence went deeper than product knowledge. An instructor had to understand how capability is actually assembled and standardized: reusable skills as packaged instruction bundles, plugins that extend a model, the way a shared library makes a team's behavior consistent.3 A teacher who cannot build a reproducible workflow cannot show a student how competence scales beyond a single clever prompt. And the instructor had to grasp that adoption is a human problem before a technical one, that the real barrier is culture and status and the fear of visible incompetence, so teaching this well means teaching change, not just tools.4
Underneath sat a pedagogical commitment the commission adopted almost without noticing: this material cannot be transmitted by lecture. It has to be practiced, learned by doing under the eye of someone who already can, the way skill is developed rather than described.5 "Which raises the obvious problem," Estrada said. "You cannot teach by demonstration a thing you cannot do. Half our proposed faculty have never operated these systems under pressure."
Then Tull did the thing that gave the chapter its edge: he drafted an instructor-qualification standard and asked the commission to apply it to itself. Several members failed. Two who had been confidently designing the curriculum could not have taught the hands-on unit they had specified; one conceded he had never verified an AI output on a live matter in his life. "This is the accreditor's recurring nightmare," Tull said, not unkindly, "and now it is ours. We have written a standard that disqualifies some of its authors." The List B beat fell out of the vendor problem: a qualified instructor had to be able to teach declining as rigorously as using, which a vendor-instructor structurally could not, because competent refusal is bad for the product. What survived was a qualification standard narrower and more honest than the commission wanted: a small pool of genuinely qualified teachers, a warning about conflicted ones, and a recorded admission that the body writing the requirements did not uniformly meet them. "We cannot mandate what we cannot staff," Klein said. "The most useful thing on this page may be the list of us who would not pass it."
Plate Thirteen
Chapter Fourteen (~2 minutes)
Assessment: What Counts as Passing
Having named the competencies, the commission had to say what a passing performance looked like, and the exercise forced a concession it had been circling since Chapter Three: some of what mattered most could not be scored, and the only honest thing was to decide, in writing, what to do about that.
Okonkwo returned to lead it, because criteria were her discipline. The first move was to define "good enough" for each kind of work, refusing the pretense that quality is binary. An internal draft, a client-facing memo, and a high-liability filing sit at different thresholds with different verification intensities, and a candidate must demonstrate she can tell them apart and calibrate accordingly.1 "A candidate who verifies a brainstorm as heavily as a brief has not shown rigor," Okonkwo said. "She has shown she cannot tell the stakes apart, which is its own failure."
The second criterion was the legal one, and here the standard was absolute. To pass, a candidate had to verify to the standard the profession actually demands, with every citation confirmed for existence, holding, currency, and relevance, and an audit trail behind it, which was the discipline whose absence had generated the entire opening docket.2 "This is not a rubric dimension we weight," Osei said. "It is a gate. You do not pass the assessment with a fabricated cite in your work product, however good the rest is."
The third criterion was the one that exposed the limit. The commission wanted to assess skill, the developed and practiced competence of doing the work well, but skill of that kind is built and shown through graded practice over time, not captured in a single sitting.3 And the analytical competence they most valued, the ability to treat a machine's output as a hypothesis and interrogate it for the buried fact and the unsupported leap, was assessable on the page but shaded, at its highest reaches, into a judgment no rubric fully caught.4 "We can score whether she flags the gap," Okonkwo said. "We cannot fully score whether she would have flagged it at 2 a.m. with a partner waiting. The best assessment in the world still leaves a residue."
The List B beat gave the chapter its resolution, because refusal turned out to be the most scorable of the judgment competencies. A candidate could be required to decline in writing, producing on a matter engineered for it a documented refusal stating the ground and the alternative, and that artifact could be graded against defensible criteria in a way that "would she have hesitated?" never could. "Declining competently is judgment you can put on paper," Klein said. "That is why it belongs in the assessment and not merely in the ideal."
What the commission decided to do about the unscoreable residue, it decided in writing, because burying it would have repeated Comment 8's sin. The assessment would certify what it could, namely calibrated quality judgment, absolute verification, documented refusal and analytical flagging, and would state on its own face the competence it did not measure: sustained judgment under real pressure, developed only in supervised practice after admission. "We are not going to pretend the passing mark means finished," Klein said. "We are going to print the limit next to the score. A profession that hides what its exam cannot measure is how we got the docket."
Plate Fourteen
Chapter Fifteen (~2 minutes)
The CLE Question
The commission arrived at continuing education expecting an easy chapter and found the most uncomfortable one, because the institution under examination was the most comfortable, and it did not enjoy the light. "The profession's current answer to all of this," Klein said, "is a CLE hour. One hour, once, to discharge a duty we have spent fourteen chapters showing is continuous and hard. We should ask, honestly, whether that hour is a credential or a fig leaf."
The case for CLE was real and Okonkwo made it: competence in this field is genuinely a continuing duty, not a thing certified once, because the tools and the standards move under the practitioner and a knowledge current in spring is stale by autumn.1 "The premise of CLE is correct," she said. "The duty never stops. The question is whether hours measure whether anyone met it."
They did not, Estrada argued, and the reason was the recurring one. Adoption and competence are human-factors problems; nearly two-thirds of organizations name culture and people, not knowledge, as the barrier, and an hour of passive attendance changes behavior in almost none of them.2 "We know how behavior actually changes," he said. "Practice, superusers, visible leadership, tiered training over months. A lecture hall satisfies a requirement and moves no one. We are measuring attendance and calling it competence."
The problem was compounded by the pace of the subject, which made the single-hour model almost comic. The field does not sit still to be certified; it moves from single-turn chat to persistent assistants to tool-using agents, each shift changing what a lawyer must understand.3 "An hour on last year's tools," Osei said, "is not merely insufficient. It can be affirmatively misleading, because it certifies a picture that has already changed." A credential that ages faster than the ink is worse than none, because it manufactures false confidence.
What a real requirement would look like, the commission sketched but could not agree on. The honest version borrowed the course's own premise, that this material is learned by doing, practiced under guidance rather than absorbed by lecture, built in graded steps that hold.4 That implied something CLE boards were not structured to deliver: assessed, hands-on, recurring practice with an actual measure of behavior change at the end. "That is not a CLE hour," Okonkwo said. "That is a competency program with an assessment. It costs real money and real time, and the boards will resist it because their entire model is seat-time."
The chapter ended, deliberately, without consensus, because the commission judged a false agreement worse than an honest split. Some members held that hours were better than nothing and reformable; others, that hours were a comfortable evasion that let the profession feel current while changing nothing, and that anything short of assessed practice was theater. Klein declined to break the tie. "I am not going to paper over this," she said. "The record should show that the body could not agree whether its own profession's main answer to this duty is a credential or a fig leaf. That disagreement is more useful to a reader than a manufactured consensus would be. Print it unresolved."
Plate Fifteen
Chapter Sixteen (~2 minutes)
What It Costs, and Who Cannot Pay
The commission had drafted a standard against the resources of a well-run firm, and Klein insisted, before adopting anything, that it be measured against the people who had none: the solo in a strip mall, the legal-aid office three lawyers deep, the rural practitioner, the small school with no technologist on staff. "If only the well-resourced can meet our standard," she said, "we have not written a competency framework. We have written an access barrier with a competency framework's manners."
The cost turned out to hide in the deployment, not the subscription. The same model reaches a lawyer through very different channels: a free public web portal, an enterprise account with retention disabled, an on-premise install, a sovereign cloud pinned to a jurisdiction. The channels that satisfy the confidentiality standard are precisely the ones that cost money and require technical staff.1 "The solo can afford the tool," Osei observed. "What the solo cannot afford is the safe way to run it. Our standard, applied honestly, prices confidentiality out of exactly the practices whose clients most need it protected."
The commission looked for the affordable path and found a real but partial one in the model spectrum. A practitioner did not need the frontier tier for everything; a fast, cheap model handles routine drafting, a reasoning model earns its cost only where legal judgment is at stake, and knowing the difference is itself the competence that keeps a small practice both safe and solvent.2 "This is where competence and access converge," Estrada said. "The lawyer who understands the spectrum spends money where it matters and nowhere else. The one who doesn't either overpays or, worse, runs high-stakes work through a tool built for speed."
But cost is not only a purchase; it is a plan over time, and the small institutions lacked the runway. Long-term implementation, meaning the sustainable practices of transparency, proportionality, and documentation that carry a practice through shifting tools and insurance expectations, assumes an ability to invest ahead of need that a legal-aid office simply does not have.3 "We keep designing for organizations that can plan," Klein said. "Most of the profession is triaging this month's caseload."
The List B beat reframed the whole chapter and, in doing so, complicated the book's own claim. For an under-resourced practice, non-use is sometimes not a principled refusal but a forced one. The tool that would genuinely help is unaffordable to run safely, so the lawyer declines because she cannot control the risk, and that is legitimate. The competence the commission could actually teach here was legal-practice application within real constraints: knowing which uses are safe on a shoestring, which require resources the practice lacks, and how to decline the latter without the client bearing the gap.4 "Refusal on grounds of cost is a real and defensible reason," Osei said, "and it is not the same as my refusal, which is a choice. Hers is a constraint. A framework that cannot tell the two apart will punish poverty and call it incompetence." What survived was a standard with a means test built into its conscience: the competencies stayed, but the commission recorded that a duty only the wealthy can discharge is an access problem the profession must solve, not an implementation detail it can assume away.
Plate Sixteen
Chapter Seventeen (~2 minutes)
The Ratchet
Before it could adopt anything, the commission had to reckon with what adoption would do, and Klein named it bluntly: the moment this framework was published, it would stop being advice and start becoming law. "Courts looking for the standard of care will find ours," she said. "We are not proposing a curriculum. We are, in effect, drafting the benchmark negligence will be measured against. That should frighten us into precision."
The mechanism was the one the commission had studied in Chapter Two and now saw from the other side. The standard of care is evolving and self-ratcheting; each articulation of what a competent lawyer should understand becomes the floor beneath the next, and a competence that is aspirational this year is presumptively mandatory once a body like theirs has written it down.1 "We looked at Comment 8 and saw a duty enforced only after harm," Tull said. "We are about to become the thing that gives the next Comment 8 its content. Whatever we publish, someone will be sanctioned against."
That raised the stakes on the framework's own risk management. If the document was going to function as a standard, it had to be built like one, on durable pillars of policy, training, quality control, and technical controls rather than a snapshot of this year's tools, because a standard pinned to specific products would be both obsolete and unfair the moment the products changed.2 "We have to write for the mechanism, not the model," Osei said. "The duty to verify survives every product. 'Use this vendor' does not, and if we write it we have built a trap."
And the subject would not hold still to be codified. The technology was mid-motion, running to persistent assistants, tool-using agents, and systems that act over hours, so any framework fixed to today's capabilities would describe a world that no longer existed by the time it bound anyone.3 This was why the commission's long-term thinking had to include its own obsolescence: sustainable standards are the ones built to be revised, flexible by design, anticipating that exclusions and expectations will shift.4 "A standard that cannot update," Estrada said, "does not stay a standard. It becomes a period piece that courts nonetheless keep applying, which is the worst of both."
So the commission did the counterintuitive thing and drafted its own sunset. Rather than trust its successors to revisit the framework, it wrote a mandatory review cycle into the instrument itself, with a default expiry that forced re-adoption on a fixed clock. There was no List B refusal beat here; the refusal in this chapter was the commission's own, a refusal to pretend its work was permanent. "We do not trust the people who come after us to update this," Klein said, "including, frankly, ourselves in five years' complacency. So we are building the ratchet with a release. The framework expires unless someone re-earns it. That is the only honest way to write a standard for a subject that will outrun us." What survived was a document that knew its own half-life: precise enough to bind, general enough to age slowly, and rigged to die on schedule unless a future commission did the work again.
Plate Seventeen
Chapter Eighteen (~2 minutes)
Commencement
The last sitting circled back to the first. Klein put the same sanctions docket on the table, and beside it a second document the commission had not had a year earlier: a framework, an assessment, and an appendix. "We opened by reading punishments for a course no one offered," she said. "We close by offering the course. Whether it was worth a year is a question the docket will answer for us, slowly, in cases we will never see."
The framework's spine was the continuing duty they had refused to let anyone discharge once, competence as a thing maintained, budgeted in real hours against tools that keep moving, never certified and shelved.1 Around it the commission had built what it could teach and examine, and had said plainly what it could not. And it had done one thing the members had not expected on the first day: it had made the case that this competence, honestly held, was not merely a shield against sanction but a professional position, since a lawyer or a firm able to show how it governs these tools holds an advantage over one that merely uses them, and demonstrable competence had become a way to keep clients and talent both.2 "We started thinking we were writing a defense against the docket," Estrada said. "We ended up describing an asset."
Two proofs closed the year. The first was pedagogical: the students who had sat the very first diagnostic, a year on, could be shown to have developed the skill in graded steps rather than merely heard about it. The ladder, climbed, produced competence you could observe.3 The second was the one Klein had insisted on from Chapter Three. The regulator who had scored worst on the commission's own audit, herself, taught a unit before the body dissolved, badly at first and then not, because the barrier had never been the technology but the willingness of a senior person to be a beginner in a room.4 "If the chair can be taught," she said, "no one in this profession gets to plead that they are too eminent to learn."
And then the divergence the commission had been honest enough never to promise it would erase. Two members would graduate committed non-users, Osei among them, and the framework did not treat them as failures or holdouts. It treated them as what Chapter Twelve had proved them to be: practitioners whose refusal was competent, reasoned, and documented, who had climbed every rung of understanding and chosen, at the top, not to rely. "The list of reasons to understand this and the list of reasons a good lawyer might decline it were never opposed," Klein said. "They were always the same syllabus read to its end. You cannot decline well what you do not understand, and you have not fully understood until you can say exactly why someone might refuse."
The appendix was the commission's last and most deliberate act: a signed public statement of what the profession still could not do, naming the judgment its exams could not score, the enforcement its offices could not fund, the access its standard did not yet reach, and the CLE question it could not resolve. It was signed by every member, including the dissenters and the two non-users, because a framework that hid its own gaps would have been Comment 8 all over again, a duty announced without the honesty to say where it fell short. "We are not qualified to administer everything we have written," Klein said, signing last. "We have said so, on the record, at the front of the document rather than leaving it for the next docket to discover. That may be the most professional thing we did all year." Then she closed the file, and the commission, having built the course it had been convened to prove was missing, adjourned.
Plate Eighteen
End · Reasonable Doubts