NEW YORK, 19 AUG 2026 — Google has won a bankruptcy auction for Spirit Airlines' internal data, bidding US$10 million against a rival offer of US$7.5 million. The package includes roughly 100 million emails, 500 million Microsoft Teams messages and 30 million lines of code, and is intended for AI training.

Every one of those emails was written by an employee who is not a party to the sale.

What was bought

US$10mWinning bid, against US$7.5m from Mercor
100m / 500mEmails and Microsoft Teams messages
30mLines of code, plus operational and pricing records
Judge approval pendingThe sale requires sign-off from a federal judge

Alongside the messages and code, the package includes spreadsheets, calendars, operational records and flight pricing history. Reported terms exclude loyalty programme members and other customer data, require the buyer not to attempt re-identification, and provide for a third party to remove personally identifiable information before Google receives anything.

The stated purpose is improving AI models and products. A federal judge must approve the sale.

Bankruptcy is a new supply channel for training data

Acquiring text to train models has become steadily harder and more expensive. Scraping invites litigation, publishers have learned to charge, and the open web has been thoroughly worked over. Against that background, a failed company's estate is an unusual asset. It is a large, coherent, privately held corpus with a motivated seller and no active business to protect.

Ten million dollars looks cheap for what changed hands, and it may still be roughly what the asset is worth. The runner-up bid US$7.5 million. Two bidders and a liquidator under time pressure is not much of a market.

This is now a category rather than an oddity. Every company that fails holds a decade of internal communication, and that corpus has just been demonstrated to have a market price. Restructuring advisers will notice. So will general counsel, who now have a fresh argument for stricter data retention policies: the only way to prevent this is to not have the messages in the first place.

What makes a corpus like this valuable

The volume is not the interesting part. Models are already trained on far more text than this.

What is scarce is real workplace communication with outcomes attached. These are not forum posts or scraped articles. They are decisions being made, escalated, reversed and explained, by people who did not know they were producing training data, alongside the operational records showing what actually happened afterwards and the code that ran the systems they were arguing about.

You cannot synthesise that, and a healthy company will not license it at any price. Only a dead one will. An AI agent designed to operate inside a business needs exactly this kind of training data. That is likely what Google is paying for.

We reported this month on the divergence in Asia-Pacific policy on training data and licensing. This transaction is the same pressure finding a route that most of those frameworks do not contemplate at all.

De-identifying conversation is harder than de-identifying a database

The reported privacy protections are better than nothing. The sale excludes customer data and loyalty records, contractually bars re-identification, and requires third-party scrubbing before delivery.

But these are controls designed for structured data, like database records. They are being applied to unstructured conversation, which is a much harder problem to solve.

Taking personally identifiable information out of a database means dropping columns. An email thread is not built that way. Free-form text identifies people constantly without ever using their names — by role, by project, by who reports to whom, by an incident only four people attended, by writing habits, by an argument that recurs across ten years of messages. You can strip names and addresses reliably. Stripping identifiability is a much harder job that goes by the same name.

The undertaking not to re-identify matters too, though it is a contractual control rather than a technical one. The contract binds what the buyer may do, but it changes nothing about the data itself. A model trained on the corpus could still reproduce identifiable patterns or text.

There is a further wrinkle specific to code. Thirty million lines of source code are not anonymous like a spreadsheet. Commit messages, comments, and variable names reveal the reasoning — and often the initials — of the developers. The systems they describe may even still be running under a different owner. Whether the scrubbing process treats code as text or as something requiring its own handling is not addressed in anything published.

The employees are the ones with no position here

The legal logic is coherent and most people will find it surprising. Work email belongs to the employer. The employer's assets belong to the estate. The estate sells assets to pay creditors. Nobody who wrote those five hundred million messages has standing to object, and nobody asked them.

The principle applies to any employee, anywhere, including you. Messages written in confidence to a colleague — about a manager, a diagnosis, a resignation being considered, a mistake — sit in a corporate system that the writer does not own and that may one day be sold by people they have never met.

For employees of multinationals across this region the exposure is the same in substance. A failed company with ASEAN operations holds correspondence written in Singapore, Manila and Kuala Lumpur, and most data protection regimes here and elsewhere stop applying once data is treated as de-identified. The protection, therefore, rests entirely on how well the de-identification process actually works.

Worth noting what the price does not cover. The auction bought a copy of a corpus, not exclusivity over what it contains, and nothing published suggests the estate is barred from selling other assets to other buyers. For a company in liquidation the incentive runs toward selling everything that will fetch a bid, and internal data has now been shown to fetch one.

What we could not establish

How the de-identification will be performed and verified. The terms mention a third-party process but do not describe it. A simple search-and-replace for names is a token gesture; scrubbing contextual details is a real protection. We do not know which this is.

Other open questions remain: Were any employees notified? How were communications from outside the US handled? What conditions might the judge impose? Does the code include licensed components? And how, specifically, will Google use the data?

What to watch

The judicial approval is the immediate thing. A court asked to sign off on a first-of-its-kind transfer may attach conditions, and those conditions would become the template for everything that follows.

Then watch whether it repeats. If a second estate sells a comparable corpus at a comparable price, this stops being an anomaly and becomes a line item in restructuring, and the interesting question becomes who else bids.

Finally, watch retention policies. The most effective response available to any organisation is to hold less for less time, and the argument for it just acquired a concrete example. Whether anyone acts on that before their own bankruptcy is a different matter.