Hackers Claim 4TB Data Theft from AI Service Provider Mercor
A supply chain attack targeting Mercor has reportedly exposed biometric data and source code from a firm serving OpenAI, Meta, and Anthropic.
Hackers have claimed to steal 4TB of data from Mercor, an AI services firm that provides training data and experts to some of the world's leading AI laboratories. The breach underscores the growing vulnerability of the specialized third-party ecosystem that supports the development of frontier AI models.
According to reports, the hacking group Lapsus$ is responsible for the theft. The stolen haul reportedly includes sensitive source code, personally identifiable information (PII), and biometric data, specifically video interviews and face scans from contractors. The breach was linked to a supply chain attack involving LiteLLM, an open-source project used to integrate various large language models.
The AI Supply Chain Risk
The modern AI industry does not operate in a vacuum; it relies on a complex network of "tier-2" providers like Mercor. These firms handle the critical, labor-intensive work of data labeling, infrastructure management, and fine-tuning. Because these providers often have deep access to the workflows and proprietary requirements of "tier-1" companies—including OpenAI, Meta, and Anthropic—they represent a high-value target for attackers.
By targeting a service provider rather than the tech giants themselves, hackers can bypass the more robust security perimeters of the primary labs. In this instance, the use of a supply chain vulnerability in an open-source tool like LiteLLM allowed attackers to gain a foothold in a company that sits at the center of the AI training pipeline.
Industry Implications
If the full extent of the 4TB breach is verified, the consequences extend beyond the loss of contractor privacy. The exposure of source code and internal workflows could provide competitors or malicious actors with insights into how leading AI labs structure their training data and optimize their models. Furthermore, the theft of biometric data is particularly severe, as face scans and video interviews cannot be rotated or changed like passwords, creating a permanent security risk for the affected individuals.
This incident highlights a systemic risk in the AI supply chain: the dependency on third-party vendors who may not maintain the same security rigor as their high-profile clients. As AI labs scale their operations, the surface area for these types of indirect attacks continues to expand.
What Remains Unconfirmed
While the involvement of Lapsus$ and the link to LiteLLM have been reported, the full scope of the data's utility to the attackers remains under investigation. It is not yet clear if any proprietary model weights or specific API keys belonging to OpenAI, Meta, or Anthropic were compromised, or if the breach was limited to Mercor's internal operations and contractor data.