Judges Apply Existing Evidence Rules to AI Deepfakes

As generative AI becomes ubiquitous, judges must update their understanding of evidence authentication.

By Central
Legal experts argue existing evidence rules remain effective against AI-generated fakes with proper judicial training.
Highlights
  • With just one minute of audio sample, software can clone a voice, making evidence fabrication easier than ever.
  • The liar’s dividend allows dishonest actors to cast doubt on authentic evidence by claiming AI generation.
  • Judicial education, not new rules, is the primary solution for handling AI-generated evidence in court.

A victim receives a series of threatening voicemail messages from a blocked number. The recordings are forwarded to police. The defendant denies making the calls and claims the audio is a deepfake imitation of his voice. The prosecutor seeks to admit the recordings as evidence. The judge must decide. This scenario, playing out in courthouses with increasing frequency, captures the central challenge facing the American legal system as generative AI becomes ubiquitous: the existing rules of evidence remain fit for purpose, but the judges tasked with applying them require a fundamentally new level of technological literacy to function as effective evidentiary gatekeepers.

The rapid evolution and accessibility of generative AI have dramatically increased the practical difficulty of the judge’s role. With as little as one minute of clear audio as a sample, commercially available software can clone a person’s voice and generate audio of that person saying anything the user types into a prompt. Never before has the fabrication of evidence been so easy and the output so convincing. Yet, this technological revolution does not diminish the continued applicability of long-standing admissibility doctrines. The most effective response today, therefore, is not the wholesale creation of new rules, but a concerted effort to ensure judges are equipped, through education and institutional support, to rigorously apply existing evidentiary principles to evidence tainted by the specter of AI generation.

The Twofold Threat to the Foundation of Justice

Our system of justice depends on the legitimacy of its fact-finding procedures. Public confidence in verdicts rests on the historic reliability of those procedures. The existing rules and doctrines of evidence, from the common law to the Federal Rules, have proven reliable over generations at separating truth from fabrication. Synthetic evidence threatens that foundation in two distinct and equally dangerous ways.

First, fabricated evidence may be credited as genuine, leading a jury to convict or acquit based on a fiction. Second, and perhaps more insidiously, genuine evidence may be successfully challenged as fabricated without any basis for the claim, giving rise to what legal scholars Robert Chesney and Danielle Keats Citron have termed the liar’s dividend. This phenomenon provides dishonest actors with plausible deniability, allowing them to discredit authentic recordings, photographs, or video simply by invoking the possibility of AI generation. Each failure carries the potential consequence of a verdict resting on something other than the truth. The cumulative risk, should such failures become frequent, is the erosion of public confidence that courts can distinguish real evidence from manufactured evidence, and ultimately, their ability to find the truth at all.

Why New Rules Are Not the Primary Answer

The natural instinct in the face of a transformative technology is to call for new laws and new evidentiary rules. However, a closer examination of legal history reveals a more pragmatic and resilient path. The introduction of DNA evidence in the late 1980s is a powerful precedent. Before it was first introduced by expert testimony in a U.S. criminal trial, a Florida court in Andrews v. State (1988) held an evidentiary hearing and considered application of a standard set in 1923: the Frye standard. The Frye standard itself was developed sixty years earlier to determine the relevancy of the systolic blood pressure deception test, a precursor to the modern polygraph. In Andrews, the court built upon Frye and affirmed relevancy as the linchpin of admissibility when considering novel evidence. The introduction of DNA did not require a bespoke rule upon which to determine admissibility. Instead, it relied on the practical evolution of then-existing principles.

This historical roadmap is directly relevant today. The Frye standard eventually gave way to Rule 702 of the Federal Rules of Evidence, with the Supreme Court in Daubert v. Merrell Dow Pharmaceuticals, Inc. (1993) further elevating the gatekeeper function of the judge. While those standards apply to scientific evidence specifically, Rule 901 applies to authentication generally. It is under Rule 901 that the judge’s gatekeeper function is tested most acutely by synthetic evidence. The law has historically applied and later adapted existing standards to new evidence types—not by rewriting the rules for every new technology, but by imposing high bars to admissibility that require an increased demonstration of reliability and authenticity, and by firmly assigning the preliminary fact determination to the judge, not the jury.

How Rule 901 Functions as a Gatekeeper for AI Evidence

Return to the hypothetical trial. The defense challenges the audio recording offered by the prosecutor. The judge, acting as gatekeeper, must decide admissibility before the jury hears it. Rule 901(b) provides a non-exhaustive and intentionally broad list of methods by which evidence may be authenticated. The most relevant in this scenario is Rule 901(b)(5), which requires an opinion identifying a person’s voice, whether heard firsthand or through a recording, based on hearing the voice at any time under circumstances that connect it with the alleged speaker. Evidence is admissible if it is relevant, meaning it tends to prove or disprove a material fact. However, subsumed in the analysis of relevancy is authentication. Evidence cannot have a tendency to make the existence of a disputed fact more or less likely if the evidence is not what its proponent claims it to be.

This is the critical doctrinal point: authentication serves as the first, and currently best, safeguard against false evidence reaching the jury. The judge evaluates whether the proponent has presented sufficient evidence to support a finding that an item is what it purports to be. In the context of a contested voice recording, a judge might look to other subsections of Rule 901. Rule 901(b)(3) permits comparison by an expert witness, such as a forensic audio analyst. Rule 901(b)(9) permits authentication through evidence of a reliable process or system, which could include metadata analysis, chain of custody documentation, or the testimony of a technical expert who can explain why the recording’s acoustic characteristics are inconsistent with known deepfake generation methods. The framework exists. The tools for its application, however, are what require urgent modernization.

The Gate Is Only as Strong as the Gatekeeper

The application of the rules is only as good as the practitioner applying them. If the judge is the gatekeeper, then the judge’s understanding of the technology is the gate itself. A gate with a broken hinge does not stop functioning; it simply stops functioning reliably, swinging open when it should hold and holding when it should open. A judge who does not understand the current capability of generative AI to create believable deepfakes cannot fulfill their obligations under the existing rules to determine the authenticity of evidence. An uninformed judge may apply the correct rules but in a way that gets the ruling wrong. The resulting error swings in both directions, and the consequences of each failure are substantial.

A judge who lacks technological literacy may allow highly persuasive but insufficiently authenticated synthetic evidence to reach the jury, effectively making twelve laypeople the first meaningful reviewers of authenticity. Once a fabricated recording is played in the jury box, the damage may already be done. Jurors are not trained to evaluate the likelihood of AI manipulation, and research, such as the study by Nils C. Köbis, Barbora Doležalová, and Ivan Soraperra published in iScience (2021), suggests that people cannot detect deepfakes but think they can. Conversely, a judge who is overly wary of deepfakes may improperly exclude authentic evidence based on unsupported allegations that it is AI-generated. This gives full effect to the liar’s dividend, allowing a defendant to discredit damning, genuine evidence with nothing more than a skeptical claim.

What Judges Need to Know: A New Baseline for Technological Literacy

In order to meaningfully apply existing rules and case law in an age of deepfakes, judges must understand the technology. They need not become experts in artificial intelligence or possess the skills of a computer scientist, but they must possess sufficient technological literacy to fulfill their gatekeeping responsibilities. This baseline of understanding includes several key components.

First, judges must recognize that generative AI can now produce highly realistic text, images, audio, video, and documents with minimal cost or technical expertise. The days when sophisticated forgery required a dedicated studio or specialized skills are over. Second, they must understand that human observation alone is no longer a reliable means of distinguishing authentic from synthetic evidence. The “seeing is believing” axiom no longer applies. Third, judges should be aware of the current limitations of deepfake identification tools. Given the trajectory of technological advances since deepfakes emerged as a recognized phenomenon in 2017, it is unlikely that detection tools will outperform creation tools anytime soon. The alliance between deepfake generation and detection is a tug-of-war, and as of 2026, the generators are pulling ahead.

With this foundational knowledge, a judge can properly assess whether the proponent has presented sufficient evidence to support a finding of authenticity. They can determine when traditional methods of authentication—such as witness testimony or chain of custody—should be supplemented by evidence of provenance, metadata, corroborating evidence, or expert testimony. Without a basic understanding of AI’s capabilities and limitations, courts risk applying established authentication principles based on outdated assumptions about the difficulty of fabricating convincing evidence.

Practical Steps: Judicial Education, Bench Books, and Technical Advisors

The solution is not a new rule, but a new commitment to judicial education. Necessary resources must be deployed to assist jurists in achieving a working standard of understanding the current state of technology. Judicial education can take many forms. Continuing judicial education programs can be developed that focus specifically on AI capabilities and their evidentiary implications. The National Center for State Courts has already begun this work, publishing resources on evaluating unacknowledged AI-generated evidence. Courts should also consider building rosters of vetted technical advisors who can assist judges in evaluating novel authentication challenges. This is a mechanism particularly needed for disputed authenticity cases that fall outside the proposed Federal Rule of Evidence 707, which currently applies only where the proponent acknowledges the evidence was AI-generated.

This approach is not a stopgap. It is a recognition that the legal system’s strength lies in its principles, not its specific rules. The Frye standard was developed for a device measuring blood pressure during questioning. It was later applied to DNA analysis. The same principle of general acceptance in the relevant scientific community remains sound; what changes is the community of knowledge the judge must consult. The Daubert standard requires the judge to assess whether the methodology underlying expert testimony is scientifically valid. In the context of AI evidence, this means a judge must be able to question an expert about training data, model architecture, error rates, and the specific conditions under which a detection tool was validated.

The Broader Implications for the Administration of Justice

The challenge of AI deepfakes is not merely a technical problem for judges; it is a systemic challenge for the entire administration of justice. Prosecutors and defense attorneys must also become technologically literate to effectively advocate for or against the admissibility of evidence. Law enforcement agencies must develop rigorous protocols for the collection and preservation of digital evidence to ensure a clean chain of custody that can withstand scrutiny. The rules of discovery may need to be applied more aggressively to require the production of metadata and provenance information. The entire ecosystem of the courtroom must adapt, but the judge remains the central figure upon whom the integrity of the process depends.

Consider, for a moment, the profound shift in the burden of proof that the liar’s dividend represents. In a pre-AI world, a party challenging the authenticity of a recording bore the burden of producing some evidence that it was altered or fabricated. Today, a defendant can simply assert that a recording could be a deepfake. The burden then shifts, in practice if not in law, to the proponent to prove it is not. This is a dramatic change. The judge must navigate this shift without explicit guidance from the rules, relying instead on their understanding of what is technologically possible and their ability to assess the credibility of the proponent’s authentication evidence.

A Roadmap for the Courts in the AI Era

The advent of AI has made the core function of a judge—to apply law to facts—more difficult, as the authenticity of those facts is now harder to verify. Questions of evidence admissibility and authenticity are not new; judges make them every day. But the tools available to help them are increasingly insufficient. The bar for authentication must rise commensurately with the ease of forgery. This does not require a new rule, but it does require a new rigor in the application of existing rules.

The historical roadmap is clear. When confronted with novel evidence, courts have consistently adapted existing standards by raising the bar for admissibility and assigning the preliminary determination to the judge. The introduction of DNA did not require a separate rule of evidence; it required courts to educate themselves on the science. The same is true for AI. The proposed Federal Rule of Evidence 707 is a useful step for cases where a party acknowledges that evidence is AI-generated, but it does not address the far more common and dangerous scenario where the origin of evidence is disputed. For those cases, the judge’s understanding is the only gate.

Return to the voicemail recordings. The judge presiding over that case does not need a new rule of evidence to decide whether they are authentic. The judge needs to understand what a smartphone and a minute of recorded speech can now produce. A judge equipped with that understanding can apply Rule 901 as it stands and reach the right result, whether that means admitting the evidence with sufficient authentication or excluding it because the proponent cannot meet their burden. A judge without that understanding will have more difficulty doing so, and the risk of error—in either direction—will be unacceptably high. Improving outcomes in the courtroom comes not from creating new evidentiary standards, but from improving judicial education and institutional resources to ensure that judges can faithfully apply existing authentication rules to evidence generated in an era of increasingly sophisticated artificial intelligence. The rules hold. The judges must be prepared to use them.

Share This Article