Flagged by an AI Classifier, Not a Hash Match: Why That's a Weaker Case.
Hash matching involves comparing the “fingerprint” of a file against the fingerprints of previously identified files. This allows for an unambiguous identification. AI classifiers, conversely, are trained to analyze image and video data and are then deployed to predict whether unfamiliar content belongs to one of a number of learned categories. The resulting AI classifier “flag” is a prediction, not an identification.
The evidentiary weight of an AI classifier flag must be assessed on an individual case-by-case basis, considering the classifier’s reliability, the context in which it was used, and the level of corroborating evidence.
As State v. Loomis, 2016 WI 68, 371 Wis. 2d 235, 881 N.W.2d 749 (Wis. 2016): “An individual must be entitled to an explanation that allows his or her legal representation to challenge the conclusions that the algorithm reached.”
Ultimately, this is one way for defendants to challenge the prosecutorial narrative: “We didn’t know what was on my computer, the AI that flags my files doesn’t make it obvious what is criminal and what is not.”
AI classifiers’ overall accuracy rates tend to hide problems with too many false positives. When targeted content is very rare, a classifier with 99% accuracy still produces lots of false flags, often far more than true identifications.
Despite the claim that the study “demonstrates that it is nearly impossible for AI classifiers to produce false positive flags,” the researchers’ data do not compare their classifier’s performance to those of hash matching software or to any baseline. The user surveys report only users’ beliefs about AI classifier results, and the researchers do not independently validate the classifiers’ results using known-good/known-bad datasets. The “real-world” examples of AI classifiers being used to exonerate or convict lack the supporting data and methodology to support the researchers’ claim, let alone evidence that AI classifiers have improved judicial outcomes.
How is a hash match different from an AI classifier flag?
Cryptographic hashes such as SHA-256 allow for an unambiguous identification by assigning a long, fixed-length number to each file. Two files with identical byte sequences will invariably share the same cryptographic hash.
Perceptual hashes, on the other hand, only identify files that are similar. Their purpose is to find similarity candidates, files that a human could reasonably conclude are the same. This requires establishing the identity of a known illegal file by human review, and then searching for similar files in other locations.
Generative AI models generate content that did not previously exist. Image classifiers, including those powered by generative AI, instead assign category scores to existing content, based on categories that the AI model has been trained to recognize.
Transcription, evidence-search, risk-assessment, and image-classification systems have different outputs and different potential error modes. None of these systems is an AI classifier.
The prompt notes that the purported reliability of AI image classifiers is “clear, given the prevalence of hallucinations in AI legal-research systems and other applications, including chatbots.” It cites an article that estimates the rate of “hallucinations” (i.e., artificial intelligence responses that are not grounded in reality or have no basis in known facts) in AI legal-research systems to be up to 10%.
But a study’s hallucination rate in one application does not establish the false-positive or false-negative rate of AI classifiers in another application. As with all forms of evidence, the reliability of the AI used against the evidence must be scrutinized in light of the specific facts at issue in each case.
What are the consequences for defendants?
The ultimate question that needs answering is whether AI classifiers’ results are as a matter of law treated as hash matches. The article the prompt references says that classifiers’ results “carry the same weight in court as hash matches for certain types of illegal content,” while noting that this appears to be a false assumption.
The prompt notes that prosecutors will call AI-generated content, AI-generated classification evidence, and AI-facilitated discoveries “the product of AI.”
The prompt concludes by noting that this means AI-facilitated evidence is likely to be used “to convince judges and juries that the evidence is more reliable, precise, and efficient than it would be if it were obtained by human efforts.”
How can an AI classifier flag an image that is legal?
An AI classifier’s confidence score is not necessarily a calibrated probability. In simple terms, a classifier’s “confidently 90% sure” does not necessarily mean there’s a 90% chance that the image or video belongs to a specified category.
If a classifier identifies an image as belonging to the category of child exploitation material, it does so on the basis of a decision threshold. Raising a classifier’s decision threshold should generally result in a lower false-positive rate, but it also increases the risk of false negatives.
As a result, the positive predictive value of a classifier’s flag is a function of the classifier’s false-positive rate and the underlying prevalence of content in the targeted category. For image content that is the subject of criminal investigation, this underlying prevalence is generally low.
In this context, even a relatively low false-positive rate can result in false flags that would substantially outweigh true identifications in a large-scale screening.
While these considerations don’t necessarily render image classifiers unreliable, they underscore the importance of evaluating classifier performance within the context of a specific content type and deployment environment. For example, even if a classifier is capable of identifying human faces, it may be incapable of determining whether the humans are underage children.
Furthermore, even when trained on huge datasets, a classifier’s aggregate performance metrics can hide issues such as poor performance with certain languages, demographics, or poor-quality images. As the European Commission recently noted in one of its publications on the Artificial Intelligence Act, “the performance of an AI system can vary depending on the type of data used for input, which can lead to biased outcomes.”
For those reasons, a classifier’s flag must be supported by corroborating evidence that makes the image’s illegal nature clear. As a result, one common defense against content-classification flags is that the classification was incorrect.
Ultimately, AI classifiers are a valuable law enforcement tool. To ensure that the technology does not lead to unjust convictions, its use must be subject to judicial review and the relevant checks and balances. This includes allowing access to the AI’s inputs, ensuring the transparency of the decision-making process, and allowing competent forensic analysis that can rebut the AI’s findings.
Can a search warrant be based on an AI classifier flag alone?
Whether an AI classifier flag alone could establish probable cause to support a warrant search depends on the circumstances and the specific AI tool.
Whether that flag is corroborated can also depend on the circumstances, and it often is. In addition to having an AI classifier flag an image, law enforcement can often rely on an image’s metadata or metadata associated with the file to corroborate a classification flag. The content of communications, information in account profiles and account activity logs, and images and videos of search queries that generated the images are also frequently corroborated sources of evidence in AI classification investigations.
In AI classification investigations, law enforcement can also often obtain information by executing a search warrant or request for information with an electronics service provider (ESP), social media platform, or content-hosting company. This can reveal corroborating evidence such as search queries, communication histories, and information about the user’s account and the content they uploaded.
In order for a court to issue a search warrant, law enforcement must demonstrate probable cause. Probable cause requires demonstrating a probable connection between suspected evidence that would prove criminal conduct and the place to be searched.
In many cases, an AI classifier flag will be corroborated by other types of evidence; and, in most cases, the corrected affidavit will still support a finding of probable cause.
Under Franks v. Delaware, a warrant issued based on a support affidavit may be invalid if the affidavit contains material omissions or falsehoods made recklessly or intentionally. Defendants entitled to a Franks hearing will show that the agent who authored the affidavit, either acted out of “reckless disregard for the truth or with intent to deceive” or made a material omission resulting in a misleading affidavit.
If a defendant succeeds at a Franks hearing, the court must determine whether the allegedly invalid warrant was supported by probable cause once the material omission or falsehood was removed or cured. This corrected affidavit is then considered, and if it still establishes probable cause, then the warrant’s validity is upheld and the evidence obtained under the warrant is not suppressed.
If, however, the corrected affidavit does not establish probable cause, then the corrected affidavit does not support the warrant’s issuance, and the warrant itself is invalid. Evidence obtained under an invalid warrant is the fruit of a poisonous tree, and defendants may be entitled to suppression.
Todd Spodek is the managing partner of Spodek Law Group, a second generation criminal defense firm that has been practicing since 1976.
Did anyone actually look at the file before police were told?
The private-search doctrine applies to many contexts. It generally holds that, where a private party discovers evidence of a crime, the government’s search is legal so long as the government’s examination does not exceed the scope of the private search.
Whether an automated referral was reliable is a question distinct from whether the government’s later search was legal. While the reliability of an automated referral is relevant to whether the government’s later search was justified by probable cause, this does not implicate the private-search doctrine.
The provisions of 18 U.S.C. § 2258A are also relevant to both the reliability of the initial reports and any subsequent government investigations.
Under Section 2258A, “covered providers” are required to report apparent violations of federal child exploitation law to the National Center for Missing and Exploited Children (NCMEC). Specifically, “a provider that obtains actual knowledge of any facts or circumstances indicating an apparent violation of section 2251, 2251A, 2252, 2252A, 2252B, or 2260 shall, as soon as reasonably possible, submit a report to the CyberTipline of the National Center for Missing & Exploited Children.”
Crucially, while Section 2258A requires reporting when applicable, it does not require the providers to monitor their users or affirmatively search for violations. Providers have reported that they are not required to affirmatively search for content that might be illegal under federal law. Instead, providers are required to “preserve such material for one year” after reporting apparent violations to NCMEC.
If ESPs, social media platforms, and other providers were required to proactively screen for CSAM, it could potentially create new liability under the Electronic Communications Privacy Act (ECPA) for accessing and disclosing users’ communications and files to law enforcement. Although some providers may choose to use AI classifiers and other means to actively screen for CSAM (and reporting by a private entity may be valid under the private-search doctrine), there is no federal law that forces covered providers to conduct such searches.
What should my lawyer demand in discovery about the AI classifier?
When fighting to challenge a classifier flag, your lawyer can demand the information that it will take to test its reliability in court. In addition to the information that is available as a matter of law, your lawyer may be able to demand the information that is available through the discovery process. This includes:
- To reproduce a historical AI classifier flag: What model version, threshold, and input did the AI classifier use to produce its output at that time? With these pieces of information, your lawyer could potentially reconstruct the classifier’s finding.
- To reproduce an AI classifier’s current output: What model version, threshold, and input does it use to produce this output? If you run the current version of the model with the same input at the same threshold, do you get the same output?
- To evaluate an AI classifier’s ability to identify illegal content without falsely flagging legal content: What is the classifier’s current false-positive and false-negative rate? A “confusion matrix” can help identify where a classifier is most reliable and where it is prone to making mistakes.
- To evaluate a classifier’s ability to identify illegal content: What validation data was used to test the AI classifier? What was the result of its test performance?
As a general rule, the “proprietary” status of an AI classifier does not establish its reliability. The mere fact that a classifier is proprietary does not entitle its owners to evade the adversarial process in court. While a protective order may restrict the disclosure of the classifier’s proprietary features to third parties, it will not protect a classifier that fails to meet the evidentiary standard of admissibility. If a defendant raises valid concerns about the AI classifier’s ability to support the plaintiff’s allegations, then the plaintiff must be prepared to address the defendant’s concerns by producing evidence that supports its claims of the AI classifier’s reliability.
Although this may seem unusual, some vendors of AI tools have yet to provide independent validation of their classifiers’ performance to the public. Instead, they offer self-assessments that serve as a testament to their products’ (supposed) efficacy. While it may be the case that these vendors’ classifiers are well-supported and highly reliable, their purported performance does not receive any public verification. The vendors’ self-interest must be considered carefully when evaluating their claims of the AI classifier’s reliability. The ABA notes: “Automated decision-making systems should be subject to oversight and accountability. It is essential to ensure that AI systems are transparent, explainable, and that the human rights of all individuals are respected.”
Can prosecutors show the jury the AI classifier's result at trial?
In federal district court cases involving image classifier flags, a common strategy for the prosecution is to use expert testimony to help explain the classifier’s output.
These claims can be challenged under Federal Rule of Evidence 702. Federal Rule of Evidence 901 requires sufficient evidence to authenticate the offered classifier record. When sufficient evidence is provided, this rule grants the defendant’s right to challenge the validity of that record. It also opens the door for the prosecution to support the validity of its record through expert testimony and the use of the record’s underlying content.
However, when this expert testimony is based on AI results from an AI classifier, it may be inadmissible. Under Rule 702, any expert testimony that lacks reliability is inadmissible.
The Judicial Conference formed a Technical Advisory Committee to help develop Rules of Evidence pertaining to artificial intelligence, machine learning, and the use of other new technologies that help create, organize, and deliver evidence. In December 2022, one of the Committee’s Judicial Conference reports proposed a new Evidence Rule 707, which would have mandated reliability review in cases where a classifier’s result was not sponsored by a corresponding human expert. If the proposal was adopted, it would have provided an additional ground for challenging the admissibility of classifier outputs at trial.
However, the proposed rule remained a proposal by the Judicial Conference.
On March 20, 2026, Congress introduced the Research and Oversight of AI in Courts Act. The proposed act requires a study on the applicability and risk of “generative AI transcription and speech recognition AI tools that are currently being tested for the Judicial Branch.” It does not specifically address the use of AI classifiers in criminal investigations and prosecutions.
Is what I typed into an AI chatbot about my case privileged?
The ruling in United States v. Heppner showed that prompts to an AI chatbot are not automatically privileged even if they concern an ongoing criminal investigation. Judge Rakoff explained that while communicating with an AI platform about legal issues does not establish an attorney-client relationship, later transmitting the information to an attorney does not make that information privileged.
The court noted that the AI platform’s privacy policy allowed it to collect and maintain logs of prompts and other data. The court also observed that these policies allowed the platform to disclose information to government law enforcement agents. The judge reasoned that users do not have a reasonable expectation of confidentiality when prompts can be collected, stored, and disclosed to law enforcement.
In the same ruling, Judge Rakoff also denied work-product protection to approximately 31 AI-generated documents that federal agents had recovered from the defendant’s home using a residential search warrant. Heppner sought protection, claiming he produced the documents in anticipation of litigation. Judge Rakoff rejected this claim, citing that not only had the defendant used an AI tool to generate them, but the resulting documents were “generic summaries of the applicable law” and “informational summaries of public information.” He characterized these documents as “AI-generated research memoranda.” Ultimately, the judge rejected Heppner’s request to suppress the documents on grounds of privilege and confidentiality.
Contact a Federal Criminal Defense Attorney
Nothing here is legal advice, and the details of your case matter. Todd Spodek and Spodek Law Group take federal criminal and white collar cases nationwide, from offices in New York, Brooklyn, Queens and Los Angeles. You can reach the firm at 212-300-5196.
Reading is good. Calling is better.
Answered within 24 hours, guaranteed. Some stories are better told out loud -
212 300 5196