Skip to main content
Can You Prove Your AI Oversight Works?Compliance Governance
5 min readFor Legal & Risk Counsel

Can You Prove Your AI Oversight Works?

The Challenge

Reaves Law Firm faced a significant issue when it submitted court documents with non-existent case citations and fabricated quotes. When defendants suspected these errors were due to generative AI, the court didn't just question the citations' validity. It demanded a detailed explanation of the firm's verification process. How did they ensure the cases existed before citing them?

The firm couldn't provide a satisfactory answer. An internal email titled "Mandatory Ethical AI Training & Reporting Protocols for All Staff" was mentioned, but there was no evidence it had been shared or that training had occurred. The general counsel had left, and filing supervision had been "restructured." Consequently, the court sanctioned Reaves under Rule 11, a standard that predates AI by decades.

This wasn't about AI producing incorrect output. It was about the firm's inability to prove that anyone had verified the output.

The Environment and Constraints

Rule 11 requires attorneys to conduct a reasonable inquiry before filing court documents. This rule applies regardless of whether you use a law clerk, a research database, or AI. The obligation remains: verify your work.

AI, however, changes the risk landscape. You can delegate research to AI quickly, and it can produce seemingly credible citations. Without a documented verification step, there's no proof that a review occurred. If the court inquires about your verification process, you're left reconstructing events from memory. If key personnel have left or processes have changed, this reconstruction becomes speculative.

Reaves operated in a setting where many law firms are experimenting with AI tools. Some have formal protocols, while others rely on general professional standards, assuming attorneys will use their judgment. The gap becomes evident when proof is required, and you realize your governance program exists only in people's minds.

The Approach Taken

Reaves tried to demonstrate oversight after the fact. The firm referenced an internal email about mandatory AI training and noted the departure of its general counsel and restructured supervision.

None of this provided evidence of verification for the specific pleadings. The court wanted to know what happened in this instance. Who reviewed the AI output? What did they check? When did the review occur? What records were created?

The firm couldn't answer these questions. It offered general assertions but no documentation linked to the specific filings. The court pointed out that the firm's namesake had signed the pleadings, and blaming departed personnel didn't absolve the firm of responsibility.

Critically, the firm ignored a warning after the first filing when defendants alleged hallucinations. Instead of reviewing its processes, the firm filed two more defective pleadings. The governance failure wasn't just about missing documentation; it was about ignoring evidence that the process was flawed.

Results and Metrics

The court sanctioned Reaves Law Firm under Rule 11. While the financial penalty wasn't specified, the reputational damage was clear. The case is now a cautionary tale in AI governance, cited in legal and compliance circles.

More importantly, the court's order highlighted the lack of a verifiable process. The firm couldn't show that anyone had performed the necessary verification. This placed Reaves in a worse position than a firm that could demonstrate a process existed, even if it wasn't followed. Records wouldn't have made the fake citations real, but they would have shown a process was in place and someone failed to follow it. Reaves faced a systems problem, not just a personnel issue.

What They Would Do Differently

Reaves would need to establish infrastructure that creates contemporaneous records of AI verification. This involves more than sending an email about training; it requires documentation tied to specific decisions.

For court filings, this might include a verification log detailing which AI tool was used, what output it produced, who reviewed it, what sources were checked, and the conclusions drawn. The log doesn't need to be complex, but it must exist and be linked to the work product.

The firm should also respond to warning signals. When the first allegations of hallucinations surfaced, that was the time to audit the process. Instead, the firm continued filing without changes. Effective governance requires monitoring and taking action when issues arise.

Finally, the firm needs to assign clear ownership. When the court asked who verified the citations, there should have been a name and a record. The departed general counsel was a convenient scapegoat, but the attorney who signed the pleadings was ultimately responsible. Ownership can't be assigned retroactively.

Takeaways for Your Team

If your organization uses AI in decision-making, consider: what would you produce if asked for proof of oversight tomorrow?

Define consequential decisions. Identify where AI influences outcomes with legal, financial, or reputational risks. These decisions need documented human review. Court filings, termination decisions based on algorithms, loan denials, and compliance determinations are all consequential.

Record the review. Ensure there's evidence that a review occurred. This doesn't mean saving every AI prompt and response, but documenting who reviewed the output, what they checked, and their conclusions. The record should be contemporaneous, not reconstructed later.

Assign ownership. A specific person should be accountable for each consequential decision. Policies stating "all staff must verify AI output" are fine as general principles, but they don't create accountability. When something goes wrong, you need to know who was responsible.

Guard against missed signals. Continuously monitor your governance. If you receive evidence that a process isn't working, stop and fix it. Reaves ignored clear signals after the first filing, compounding the governance failure.

Your AI policy may promise human oversight, but that's just an assertion. The real question is whether you can prove compliance when asked. If your answer relies on reconstructing events from memory or pointing to departed employees, you don't have a governance program. You have a document.

Rule 11

You Might Also Like