BIPA

When Does Public Audio Training Become Voiceprint Collection?

When Does Public Audio Training Become Voiceprint Collection?

Google recently asked the U.S. District Court for the Northern District of Illinois to dismiss a proposed class action involving AI and voiceprints. In the complaint, a group of journalists, voice actors, and audiobook narrators allege that Google used their publicly available recordings to train voice models without notice or written consent, collecting voiceprints protected by the Illinois Biometric Information Privacy Act (BIPA) in the process.

Google argues that the complaint does not specify which recordings actually entered its training data and does not allege non-conclusory facts showing that the company extracted the plaintiffs’ voiceprints. The complaint states that Google’s products can generate natural-sounding speech, but it does not allege that the resulting AI voices would be recognized as belonging to any plaintiff. Meta has raised similar arguments in a related case.

Google’s request for dismissal indicates that the inclusion of public recordings in training data is not, by itself, enough to establish the collection of a voiceprint under BIPA. The analysis requires closer consideration of BIPA’s provisions, the parties’ arguments on dismissal, and existing case law.

The court has not ruled that Google violated BIPA. The allegations concerning the ingestion of recordings and extraction of voiceprints remain the plaintiffs’ claims, while the positions advanced by Google and Meta are defenses raised in their motions to dismiss.

I. BIPA’s Definition of “Voiceprint”

BIPA took effect in 2008 as an Illinois law specifically governing biometric data. Section 10 lists a voiceprint, together with a retina or iris scan, fingerprint, and scan of hand or face geometry, as a “biometric identifier.” Under BIPA, “biometric information” means information based on a biometric identifier and used to identify an individual. The statute does not further define “voiceprint.”

Section 15 establishes several obligations for handling biometric data. A company in possession of such data must develop and publish a retention and permanent-destruction policy. Before collecting, obtaining, or purchasing the data, it must provide written notice of the collection and storage, their purpose and duration, and obtain a written release. The statute also restricts sale, profit, and disclosure and imposes data-security requirements. Section 20 authorizes an aggrieved person to bring an action and provides remedies including statutory damages, attorneys’ fees, and injunctive relief.

BIPA identifies the biometric identifiers it protects and the duties that apply to them, but it does not state when ordinary processing of a recording becomes voiceprint collection. This gap in the definition has become a central source of dispute in practice.

II. Google and Meta’s Main Grounds for Dismissal

The plaintiffs identify publicly released recordings of their professional work and describe Google’s development of voice synthesis, automatic dubbing, and other AI voice products. Publicly available information does not establish whether particular recordings entered Google’s training data, so the plaintiffs make the relevant allegations on “information and belief.”

Google’s request for dismissal focuses on three points. First, the presence of recordings on the internet does not show that Google actually used them. Second, a product’s ability to generate natural-sounding speech after training does not show that Google created or stored voiceprints for the plaintiffs. Third, the complaint does not allege that Google’s AI voices would be recognized as belonging to any plaintiff or offer other specific facts showing that the relevant representations are capable of identifying an individual.

Meta’s defense in the related case also centers on data ingestion. Meta argues that the plaintiffs do not allege facts showing that the company ingested their specific recordings and that their claims rest on speculation.

At the motion-to-dismiss stage, the court asks whether the complaint alleges facts sufficient to support a reasonable inference. It does not decide whether the plaintiffs can ultimately prove every claim. If the cases proceed to discovery, training-data inventories, dataset sources, feature-extraction methods, and model-evaluation capabilities may become central points of dispute. How courts address information held primarily within the companies will also affect the pleading threshold in similar AI training-data litigation.

III. Relevant Case Law

Existing case law has not directly answered when using public recordings for AI training becomes voiceprint collection. Carpenter concerned the processing of customer voices by a drive-through voice system, while Delgado concerned speech that users submitted directly to Facebook and Messenger. The cases involve different factual structures, but their rulings examine the roles of general vocal characteristics, technical representations, and identifiability in a BIPA analysis.

A. Carpenter v. McDonald’s

Carpenter v. McDonald’s Corporation concerned an AI voice assistant used in McDonald’s drive-through restaurants. The plaintiff alleged that the system mechanically analyzed customers’ voices and combined the results with other information to identify repeat customers. McDonald’s maintained that the system only interpreted what customers said and did not extract biometric data.

In its 2022 ruling on the motion to dismiss, the court stated that characteristics such as pitch, volume, duration, accent, and speech pattern, viewed individually, are insufficient to identify a person uniquely. The complaint also alleged, however, that the system used an acoustic model to take mechanical measurements and used the resulting data to identify unique customers. The court held that these allegations supported a reasonable inference of voiceprint collection at the pleading stage, while emphasizing that the facts remained far from proven.

Carpenter shows that the distinction between general vocal characteristics and a voiceprint depends on the combination of features, the technical measurements taken, and their identification function. A system’s ability to interpret speech or produce a transcript does not independently establish voiceprint collection. Using the system to recognize the same speaker across different settings provides more direct evidence for that conclusion.

B. Delgado v. Meta

Delgado v. Meta Platforms, Inc. concerned speech that users submitted directly through Facebook and Messenger. Citing Meta’s voiceprint patents, privacy notices, and product-processing methods, the plaintiff alleged that Meta could create voiceprints from speech and use them to identify users.

In its 2024 ruling on Meta’s motion to dismiss, the court held that the plaintiff did not need to prove at that stage that Meta had actually used a voiceprint to identify her. The complaint cited patents describing methods for creating digital voiceprints from audio and using them to identify or verify users. Meta’s privacy notice also stated that it might collect recordings that could be used to identify users. Taken together, the court found these materials sufficient for the voiceprint claims to proceed.

After the case entered discovery, Meta moved for summary judgment. It argued that it had collected only voice recordings and that the record did not show that the company created output representations capable of identifying the plaintiff. The court denied the motion in May 2026. It found that the expert opinions and system records created a genuine dispute of material fact over whether Meta processed the plaintiff’s voice in a manner capable of identifying her. BIPA focuses on whether the data is capable of identification; the plaintiff does not need to show that Meta actually completed an identification. The court also explained that its ruling did not precisely define the point at which an ordinary recording becomes a voiceprint.

Delgado remains factually distinct from the Google case. The plaintiff in Delgado used Meta’s voice functions directly and connected the product to identification capabilities through patents and privacy notices. The Google case concerns professional recordings distributed through public channels, and the plaintiffs must still connect those recordings to specific training data. Delgado’s analysis of identification capability cannot substitute for facts showing data ingestion in the Google case.

IV. Factors for Determining Voiceprint Collection in Public-Audio Training

Drawing on Google and Meta’s arguments for dismissal and the analysis of vocal characteristics and identifiability in Carpenter and Delgado, the question of voiceprint collection in public-audio training can be organized around three related factual issues. BIPA does not list these issues as independent statutory elements. This article uses them as an analytical path derived from the current disputes and case law.

A. Whether Specific Recordings Entered the Training Data

Google’s and Meta’s dismissal arguments begin with data ingestion. In addition to showing that relevant recordings were publicly available, plaintiffs need to allege facts supporting the conclusion that the defendant actually used those recordings. The large volume of interviews, voice-over work, and audiobooks available online shows only that a company could have accessed the material. A model’s ability to generate natural-sounding speech also does not establish that a particular plaintiff’s recordings entered the training set.

Training-data inventories, procurement or licensing records, scraping logs, dataset versions, and training-task configurations can show whether specific material was ingested. These records are generally held by model developers or data suppliers, so plaintiffs may make some allegations on “information and belief.” The court must still decide whether the public facts identified in the complaint move data ingestion from a possibility to a reasonable inference.

B. Whether Training Created a Technical Representation Linked to an Individual

After a recording enters a training set, a model may extract pitch, timbre, rhythm, accent, volume, pronunciation, or other acoustic features. A BIPA analysis must then consider whether those features were combined into a technical representation linked to a particular individual. Carpenter indicates that an individual vocal feature ordinarily cannot identify a person uniquely and does not automatically become a voiceprint merely because it has been numerically encoded.

The same recording can support different technical uses. Speech transcription focuses on what was said. Speech synthesis learns pronunciation, prosody, and linguistic patterns. Speaker identification and verification focus on the speaker’s identity. The analysis therefore turns on how the company combines and uses vocal features. Whether the training process creates a speaker embedding or identity label, whether the representation remains linked to the source speaker, and whether the model supports cross-recording retrieval, voice verification, or imitation of a specific person can all affect the voiceprint analysis. Model parameters used to learn general linguistic or acoustic patterns present a materially different factual structure from a stored voice template created for a particular person.

C. Whether the Representation Can Identify an Individual

Both Carpenter and Delgado treat identifiability as an important fact in the voiceprint analysis. At the motion-to-dismiss stage, a plaintiff may not need to prove that a company has used the relevant data to identify them, but the complaint must allege facts sufficient to show that the system is capable of doing so. Delgado’s 2026 ruling further indicates that internal systems capable of processing vocal features and linking recordings to account data can create a genuine dispute of material fact over identifiability.

Evidence of this capability may appear in product functions, training objectives, patents, privacy notices, identity labels, speaker-verification tests, similarity thresholds, and internal evaluation materials. Labels such as “voice model,” “acoustic features,” or “embedding” do not determine the legal characterization. Realistic AI output primarily reflects synthesis quality, while an output that resembles a professional voice type or accent primarily reflects stylistic similarity. Those facts require an additional connection to individual identification before they can support a claim of voiceprint collection.

Together, the three factual issues connect the source of a recording, the processing performed during training, and the identification function. For companies that train voice models on public recordings, records documenting audio provenance, feature formation, and model-identification capabilities bear directly on whether Section 15 obligations apply. In litigation, those records may also supply or break the links in the asserted factual chain.

The Google case remains at the motion-to-dismiss stage. The court has not determined whether the recordings entered the training data or whether Google collected the plaintiffs’ voiceprints. The next ruling will assess whether the facts currently alleged in the complaint satisfy the pleading threshold. If the case proceeds, discovery may provide a basis for examining training-data sources and model-processing methods more closely. The extent to which the case informs other uses of public recordings for AI training will continue to depend on the data sources, processing methods, and identification capabilities the court ultimately finds.

Sources

Illinois Biometric Information Privacy Act, Section 10; Illinois Biometric Information Privacy Act, Section 15; Illinois Biometric Information Privacy Act, Section 20; Complaint in Marin et al. v. Alphabet Inc. et al.; Ruling in Carpenter v. McDonald’s Corporation; 2024 motion-to-dismiss ruling in Delgado v. Meta Platforms, Inc.; 2026 summary-judgment ruling in Delgado v. Meta Platforms, Inc.; American Bar Association: Voiceprints, AI, and BIPA; MediaPost: Google, Meta Seek Dismissal of Voice Actors’ Suits; Law360: Meta Says AI Voice Suit Rests on Speculation, Not Facts.

Start Your Compliance Journey !

Contact security and privacy veterans at Kaamel

https://kaamel.com
info@kaamel.com
340 E Middlefield Rd, Mountain View, CA 94043
AICPA Drata
© 2024 Kaamel Inc. All rights reserved.