AI and "It depends"
What a fabricated airline case and a secret risk score reveal about responsibility in law

Air travel often presents its challenges, from cramped seating to harried cabin crew navigating crowded aisles. Its not rare for us to wait for a delayed flight or suffer under a noisy travel companion making us wonder if we can sue the airline. In August 2019, Roberto Mata embarked on an Avianca flight from El Salvador to New York. His journey took an unexpected turn when, he alleged, a metal serving cart struck his knee. Such claims are not uncommon, as airlines frequently face lawsuits, this incident initially appeared unremarkable.
Mata's attorney, Peter, turned to ChatGPT for supporting legal precedents and was provided with five non-existent cases. When questioned about the authenticity of one of these cases, the chatbot erroneously confirmed its validity. Consequently, in June 2023, Judge P. Kevin Castel imposed a $5,000 penalty on the lawyers and their firm, holding them jointly and severally liable. Separately, Mata's original claim was dismissed due to being time-barred.
The machine's error became consequential due to the lawyers failure in verifying its output. Relying on the source of an invention to confirm its own accuracy does not constitute an independent or reliable check.
When confronted with a question lacking a definitive answer, lawyers often resort to two crucial words: "it depends." Properly employed, this phrase precisely identifies the variables influencing advice, such as missing facts, unsettled legal rules, or risks requiring deliberation. While AI models can also generate these words, discerning whether their 'hesitation' reflects genuine uncertainty or merely simulated caution remains a distinct inquiry.

A different kind of failure emerged in State v Loomis. Eric Loomis's pre-sentence report featured a score from COMPAS, a proprietary tool designed to assess the risk of reoffending. The sentencing court referenced its high-risk assessment when imposing six years of confinement and five years of extended supervision. Loomis challenged this score, particularly because its weighting algorithm was protected as a trade secret. In 2016, the Wisconsin Supreme Court allowed the use of such tools, but with specific restrictions: a COMPAS score could not singularly determine imprisonment or sentence severity, and all reports had to include warnings about the tool's inherent limitations.
Unlike a chatbot, COMPAS functioned as a predictive instrument; its fundamental flaw lay in opacity rather than outright invention. In Mata's case, the legal reasoning was fabricated, whereas in Loomis's case, it was simply unavailable for scrutiny. While the legal authorities cited in Mata could have been independently verified, the proprietary weighting of COMPAS in Loomis severely restricted any independent examination.
This challenge of unreliable AI persists within the legal field. In June 2025, the Divisional Court in England and Wales confronted the issue of false authorities in Ayinde and Al-Haroun. The court underscored lawyers' critical duties of verification and highlighted the potential for regulatory referrals and, in appropriate circumstances, contempt proceedings. Even specialist commercial legal research products are susceptible to these errors; a 2024 study revealed hallucination rates ranging from 17 to 33 percent across the systems tested.
While these advancements do not inherently render the technology a liability, the true value of AI in tasks like document review, chronology generation, clause comparison, and drafting initial documents is directly proportional to the extent of human verification required for its output. Highlighting this need, Abu Dhabi's Judicial AI Platform, launched its first phase in September 2026, explicitly emphasizes human oversight and verification as central to its operation.
Regulators have sought to delineate AI applications by task: for instance, the EU's AI Act designates systems assisting judicial authorities in interpreting facts and law as high-risk, whereas purely administrative work that does not impact individual cases falls outside this classification. While useful, this distinction remains incomplete; even a seemingly minor tool, such as a summary, can significantly influence a judge's perception of a file, and a retrieved legal authority can mislead if its relevance is misinterpreted.
Even as AI systems improve in reliability, and acknowledging that human lawyers are not infallible, the fundamental practical questions persist: Can the provided authority be independently verified? Does the conclusion logically derive from the presented material? And critically, is the human reviewer empowered and able to challenge the AI's output?
Simply labeling an AI system as an 'assistant' does not resolve these critical questions. Effective human oversight demands an individual who is not only willing but also adequately equipped to challenge and reject AI-generated output. After all, a signature affixed to an unread draft provides no meaningful protection to anyone.
Ultimately, the extent to which AI can be integrated into legal practice hinges on several factors: the specific tool employed, the nature of the task, and the potential cost of error. Regardless of AI's involvement, the individual or entity relying on the AI-generated output remains ultimately accountable for its correctness and consequences.
Goutham Krishna is a Bahrain-based legal specialist, entrepreneur, and media owner with established expertise in corporate law, government relations, and regulatory compliance. He was called to the Bar at Lincoln's Inn and he holds advanced qualifications including an LLM in Public Law and an LLB from Queen Mary University of London and a degree in Economics from Loyola College, Chennai. Currently, Goutham serves as Founder & Legal Director of Noor AL Qallam, and as Legal Director of News of Bahrain.