Google buys Spirit Airlines corporate data for $10M to train AI
The tech giant won a bankruptcy auction for internal communications and operational records to improve model reasoning.
Google has acquired a massive trove of internal business and operational data from the bankrupt budget carrier Spirit Airlines for $10 million. The purchase, finalized through a bankruptcy auction, provides Alphabet with a vast array of corporate communications, financial records, and passenger data to train its artificial intelligence models and support product development.
Google secured the dataset by beating a competing bid of $7.5 million submitted by Mercor, an AI data startup. The acquisition includes the airline's internal software code as well as corporate communications, including emails and chats. To comply with legal requirements, the bankruptcy court has mandated that all data be deidentified before it is transferred to Alphabet.
The hunt for corporate data
Spirit Airlines entered bankruptcy proceedings, which led to the auction of its non-core assets. This sale comes as AI developers face a critical shortage of high-quality, real-world corporate data to train Large Language Models (LLMs). While public web data has been the primary source for AI training, internal corporate archives—often referred to as "dark data"—are now highly prized. These records, which include operational logs and professional communications, provide the nuanced, domain-specific knowledge necessary to improve a model's reasoning capabilities in professional environments.
Industry implications
This deal underscores the aggressive pursuit of private corporate archives by tech giants seeking a competitive edge in AI sophistication. By integrating the operational history of a major airline, Google can better train its models on complex logistics, corporate decision-making, and industry-specific workflows. However, the transaction also raises significant ethical and privacy concerns. The sale of employee communications and passenger records during a bankruptcy liquidation highlights a growing tension between the commercial value of corporate data and the privacy expectations of the individuals who generated that data.
What remains
As Google begins integrating this dataset into its training pipelines, the industry will be watching for how this specific corporate knowledge manifests in future AI product updates. While the scale of the data is confirmed as massive, the exact volume of records remains a point of discussion. Observers will also be monitoring whether this sets a precedent for other bankrupt corporations to liquidate their internal communications as AI training assets.