Medical Photos Flagged as CSAM: The Case That Proved the System Fails.
In February 2021, Google flagged images sent from two fathers’ accounts to a doctor, whom they were treating their children remotely via photo consults. The photographs, for which a doctor had specifically requested them, showed genital infections.
The report filed with the National Center for Missing & Exploited Children (NCMEC) cited “an automated image analysis” as the trigger, not a hash match.
After obtaining access to the imagery, police found they clearly bore the markers of medical photographs. They filed no charges against the fathers.
While agencies like the FBI and DOJ have updated the language to “CSAM,” federal statutes still use the term “child pornography.” For example:
- 18 U.S.C. § 2256(a)(4)(C) defines “covered depiction” to include “a visual depiction that . .. appears to be a minor engaged in sexually explicit conduct,”
- 18 U.S.C. § 1466A(b) applies to “child pornography” that is “obscene,” regardless of whether it depicted “an actual minor.”
However, the Supreme Court in Ashcroft v. Free Speech Coalition (2002) limited bans that apply to virtual child pornography because, as it explained, such speech “records no crime and creates no victims by its production.”
So, why did the detection system do it?
It was a failure of the system.
The evidence was clear: the detection pipeline’s automated triggers are unable to discern context. The automated image analysis failed to distinguish legitimate medical imagery from CSAM. The hash detection system does not distinguish between legitimate medical imagery and CSAM either; it simply flags files that match the hash of a CSAM image in its database. This happened even though the images were sent to a legitimate doctor. Both of these are examples of the detection pipeline convicting (or at least flagging) context it cannot read.
The law recognizes this as well, as seen in Ashcroft v. Free Speech Coalition. The Supreme Court wrote:
What Happened After Google Detected the Fathers’ Medical Photos?
These two cases arose within one day of each other, with same-day images being flagged by Google’s automated image analysis system. In both cases, Google employees reviewed the photographs, saw that they were of medical infections, and upheld their “CSAM” classifications. Then, Google reported both fathers to the National Center for Missing and Exploited Children (NCMEC). These reports, or “CyberTips,” are the formal reports that law enforcement use to track down the “perpetrators” and “victims” of CSAM and other forms of abuse. NCMEC then reviews these CyberTips and makes them available directly to the appropriate local, state, or federal law enforcement agencies.
NCMEC states that a CyberTipline report is an investigative lead rather than a confirmed finding, and reports may be generated by hash matches, automated classifiers, or human review. As a result, law enforcement is able to see a report of abuse from NCMEC and assume that the report comes from a known CSAM image.
In both cases, the reporting process proceeded to law enforcement and the police obtained access to the imagery. However, these are cases that the police closed immediately. After seeing the photos and talking to Mark and Cassio, police officers in San Francisco and Houston recognized that the photographs bore the markers of medical photography, and filed no charges against Mark or Tom.
Neither father was arrested, prosecuted, or convicted over the doctor-requested photographs. However, as a result of having them flagged as CSAM, Mark lost his Google account. After the police had obtained access to the imagery, they reached out to Google to tell them how to move forward. Then, the police told the fathers that they would not be filing charges. But Google kept Mark’s account disabled and permanently locked out.
While Google had designated Mark as the “perpetrator” in a probable case of child sexual abuse, Mark never had the chance to defend himself. To open a case, Google, NCMEC, NLEC, or police must find an image of a child being abused. In the cases of Mark and Tom, no child was being abused. No one wanted to be able to open a case against them. Google’s detection pipeline simply decided they were “perpetrators” in a “documented hash match.”
While Mark is not facing any consequences for CSAM or any criminal charges for sending medical photos, he is still facing consequences for being identified as a CSAM perpetrator. His account remains disabled. Mark lost access to his Gmail and Google Photos accounts. As result, he has lost access to all of his emails and all of his family photographs. Because his account was also the one he used to purchase the phone he now has, Google disabled his phone as well, and the phone became a paperweight.
To make matters even worse, Google’s decision to leave the account disabled remained a separate moderation decision from the police’s decision not to prosecute. After obtaining access to the photographs, law enforcement was able to see that they were medical in nature. They were able to see that no charges should be filed. And no charges were filed. But Google continued to deny Mark the use of his account and phone.
CyberTips typically include:
- Account identifiers
- Timestamps
- IP addresses
- Filenames
- Hash values
Was Google Using Hash Matching or Artificial Intelligence?
What is a Cryptographic Hash?
A cryptographic hash creates a “fingerprint” of a file based on the file’s underlying bytes. If the image has been converted into a different file format, if it has been cropped, if the resolution has been changed, or even if the user just cropped a few pixels, the cryptographic hash will be completely different. If two files have a hash match, this means that the two files are identical. Even if they have different filenames or were uploaded at different times by different users, they contain the exact same image.
How is This Different From a Hash Match?
It is not different. It is a specific type of hash match.
What is a Hash of a Known CSAM File?
If a file’s hash matches the hash of a known CSAM file, that identifies the file as containing CSAM. This does not identify the user, though, so the reporter must provide the user’s ID, IP address, or other identifiers.
What is PhotoDNA?
PhotoDNA is a perceptual hashing system, and other perceptual hashing systems exist. Unlike a cryptographic hash, a perceptual hash recognizes an image even if a version of it is slightly different from the version in its known-CSAM database (for example, the resolution is different or it is in a different file format). Similarly to the cryptographic hash, a perceptual hash identifies content rather than users.
How Does This Differ from a Hash Match?
Again, it is not different. It is a specific type of hash match.
What is an AI Image Classifier?
An AI image classifier is an artificial intelligence tool that can identify CSAM and other imagery. An AI classifier is different from a perceptual or cryptographic hash because an AI classifier will be able to identify CSAM imagery even without a perceptual or cryptographic hash match in its known-CSAM database. However, AI classifiers that are used for these purposes will generally produce a perceptual hash as a part of their classification process.
How Are These Different?
Cryptographic hashes, perceptual hashes, and classifiers all look the same at a basic level, but they are quite different when compared in detail. Cryptographic hashes compare binary data. Perceptual hashes (such as PhotoDNA) compare visual data. Classifiers use more complex algorithms to identify features or object patterns within the image data.
What Are the Implications for Reporting?
Similar to this difference in how the systems recognize imagery, these are also different in terms of the evidentiary records that they generate. While a user report, an automated classifier, and a hash match might all generate a CyberTip, a record of an AI classifier’s identification is one type of evidentiary record, and a record of a cryptographic hash’s identification is another. This means that these reports can be distinguished from one another.
What about a “Hash Collision” with a Cryptographic Hash?
A hash collision is when two files that aren’t identical share the same cryptographic hash, which means that a detection system can mistakenly identify one file as another. While this is theoretically possible, it is extremely rare, and it is not considered a routine explanation for a match.
What about the Possibility of a Perceptual Hash Collision?
A perceptual hash collision occurs when the perceptual hash algorithm wrongly recognizes two visually distinct images as identical. But again, there is no evidence of this type of collision in either of these two cases.
Todd Spodek and the attorneys at Spodek Law Group handle federal cases of this kind from New York, Brooklyn, Queens and Los Angeles.
Why Did Human Review Still Produce a CyberTip?
These medical-photo cases help explain why the detection pipeline continues to fail.
As discussed above, the problem is not that the detection systems are unable to understand imagery, but rather that they are unable to understand context. The medical-photo cases demonstrate that human review also has the potential to miss external context. The only reason that the image content alone was not enough to tell the human reviewers the correct context was because of how images themselves function. Without access to the surrounding context (like the conversations that led to the images being sent), images simply are not the kind of evidence that humans are able to use to deduce external context in the same way that AI tools can be trained to deduce internal context.
As a result, if a provider finds that any image falls within the “child sexual abuse imagery” or “child pornography” thresholds, the provider is likely to proceed with the reporting process regardless of the context of the image or images in question.
When a provider determines that an image (or a set of images) is CSAM, the provider is required to send a qualifying report to the National Center for Missing and Exploited Children. Once the provider sends the report through NCMEC’s online reporting tool, the CyberTipline, NCMEC forwards the reports to the appropriate local, state, or federal law enforcement agencies. Because the investigators receive a CyberTipline report, not a concluded investigation, the report is intended to be used as an investigative lead, rather than as conclusive proof.
As long as Google has determined that the medical photographs of the children’s infections “were of CSAM content,” it will be considered a case of CSAM and will be reported accordingly. This means that a “review threshold” does not supply any additional medical or conversational context. Instead, it relies on image-level context, which, as the medical-photo cases show, is insufficient for providing an accurate diagnosis.
Law enforcement officers (and prosecutors) are generally familiar with the difference between various thresholds. These are most famously applied in criminal law:
- “Mandatory reporting” implies that a crime was committed and that the reporter has a legal obligation to report it.
- “Probable cause” suggests that a crime was committed and that there is a “reasonable probability” that the reported suspect committed the crime.
- “Beyond reasonable doubt” is a highly rigorous standard used in criminal trials to ensure that there is a high probability that a person committed the reported crime.
However, these distinctions aren’t applied to the laws that mandate the reporting of CSAM. For example, the law that mandates reporting for the various technology companies and service providers under 18 U.S.C. § 2258A says:
“A provider . .. that has actual knowledge of any facts or circumstances from which there is an apparent violation of section 2251, 2251A, 2252, 2252A, 2252B, or 2260 shall not knowingly . . . .. and . .. shall, as soon as reasonably possible after obtaining actual knowledge, submit a report to the CyberTipline of the National Center for Missing & Exploited Children in the manner and format that the Center prescribes.”
Under this law, the provider is required to report if they have “actual knowledge of a violation that appears to be in violation” of these laws. This is far from a case of “proof beyond a reasonable doubt.” In the two cases described above, the providers’ “knowledge” of the violations was based on automated detection and a lack of medical or conversational context from the human reviewers. As a result, the law triggered the reporting of the image and identified the fathers as suspects in CSAM cases.
Can Police Make Google Restore a Cleared Parent’s Account?
Do Platforms Have to Restore Accounts Based on Police Requests?
Not necessarily. The reason is that platforms have separate moderation policies for banning users and accounts. This moderation does not have to meet the standard of proof that is required to charge a suspect under criminal law.
In the cases described above, reports were first flagged by the platforms. The platforms then forwarded the reports to the CyberTipline. Google then referred each suspect’s case to local law enforcement agencies. The referral did not shift control to the police. The platforms themselves have account control, including the power to disable accounts or to restore access after a case is settled.
Can the Providers Restore the Account?
Yes. As long as it is not prohibited by a contract or court order, the providers have full control over whether or not to restore the account after a case is cleared. As it is most frequently the case, providers will decide when and how to restore access to their services once they decide the account holder is no longer in violation of their platform’s moderation policy. This means the decision is still a private moderation decision, not a government mandate.
What Are the Consequences of This Lack of Control?
The problem is that the account is disabled before any judicial review takes place. This leaves the accused suspect with absolutely no way of regaining access to their account without a formal request from the provider. For instance, the suspension of the account removes access to the user’s email before a criminal court adjudicates it. This interferes with the user’s ability to use not only their computer and their phone but also the account itself and its connected services. If their account was tied to other services, then they will lose access to two-factor authentication as well, since they no longer have access to their email account or their phone.
A result of the provider’s moderation decision to disable the user’s account is also that the user cannot access evidence stored in their cloud account. And while the police have access to the files, the provider is not required to give access to the files back to the user. This is further complicated by Section 504 of the Adam Walsh Act. Section 504 prohibits any duplication of the suspect’s account content once identified as “CSAM.” For example, if an image has been flagged as “CSAM,” then anyone who downloads or even stores the suspected CSAM file could be criminally liable as well. As a result, the law requires that CSAM evidence remain in the custody of the government. This allows the police to avoid duplication of the content and, in turn, does not require that copies of any suspected CSAM content be returned to the suspected “perpetrator.”
How Can Parents Document Telemedicine Photos Before Sending Them?
The medical-photo cases show how risky it is for parents to have medical-purpose photographs in their digital life. However, it is possible to document the medical nature of the files before the image is captured. One way is through a clinician’s written request. For example, if a doctor tells a parent, “Take photos of the rash on your child’s thigh,” a parent can request a written copy of the instructions in order to be able to prove the medical purpose of the image or images later. Another way is to use the clinician’s secure medical portal, since this allows sensitive photographs to stay inside the patient’s medical communication record.
Are There Any Risks of Using a Medical Portal?
Yes. When a parent uses the camera on a smartphone to take an image, the phone often defaults to the device’s camera roll. From there, the phone will automatically copy and upload the image to a cloud provider. Once the image is uploaded to a cloud provider (even if the provider is not the clinician’s medical portal), it will trigger the detection pipeline discussed above.
What is the Best Way to Take the Photo to Avoid This Risk?
Many smartphones allow users to take photographs directly within an application. This means that parents can use their smartphones’ cameras without copying the images to their photo backup service. Once the photo is uploaded to the clinician’s medical portal through their smartphone’s camera application, the photo is kept out of the cloud provider’s automatic cloud-backup system.
How Can I Document the Medical Nature of My Images if I Already Sent Them?
You can still preserve evidence of the medical nature of the images if you have already sent them. You can preserve the relevant chat log and appointment record that show the reason for sending the images. In a criminal case against your parents, for example, the records of a chat between the parents and a doctor that requested the image(s) could help establish a lack of intent and a lack of knowledge of the images being CSAM or child pornography. To further mitigate risks, parents who need to send sensitive images over the Internet should consider disabling their automatic cloud backup service before they send images to an approved clinician.
How Can I Document the Telemedicine Context by Using Evidence from the Device?
If a defendant is facing charges in a telemedicine case, defense attorneys can also seek to recover relevant-related-information. This means a forensics expert can look at the evidence and preserve the records of biometric unlock, app-usage, Windows Jump Lists, shortcut files, and the shellbags. These sources can help determine if the user opened the image(s) in question, used certain applications (such as the messaging application they used to send the medical image to the doctor), or accessed relevant folders.
Get Advice on Your Situation
If you want someone to look at the specifics of your case, Spodek Law Group handles federal criminal defense nationwide from New York and Los Angeles. The firm has been practicing since 1976 and its motto is simple: we owe loyalty to only you. Call 212-300-5196.
Reading is good. Calling is better.
Answered within 24 hours, guaranteed. Some stories are better told out loud -
212 300 5196