Privacy & Data 6 min read

OpenAI Has Hundreds of Contractors Reading Whole ChatGPT Conversations

They see whole conversations and score them 1 to 7. Names are stripped; the memory summaries describing where you live and what you do are not.

Priya Nair
Data, AI Governance & Policy Analyst
Published 16 Sep 2026, 4:22 PM (SGT)
Share:
Two hands typing on a black laptop keyboard, photographed from above. Two hands typing on a black laptop keyboard, photographed from above. Photo by Szabó Viktor on Pexels
Advertisement

16 SEP 2026 — OpenAI has hired hundreds of contractors to read real ChatGPT conversations and grade the answers, according to internal material obtained by 404 Media. The programme is called Project Lily, and the readers are not OpenAI employees.

Contractors are recruited through a firm called Crossing Hurdles and paid through Mercor, with one worker reporting more than $50 an hour. They see whole conversations, not isolated messages.

What a reviewer actually does

The task is not moderation. Reviewers summarise what the user was trying to do, judge how the model responded, mark the passages that meet or miss the requirement, and score the answer from 1 to 7, where 1 is unusable and 7 is close to ideal. They write a justification for the score.

The instructions shape the model's manner as well as its accuracy. 404 Media reports that the material includes training ChatGPT not to anthropomorphise itself and to be less sycophantic, which is a reasonable description of two complaints users have made loudly for a year.

This is ordinary machine-learning practice. Models learn to choose between two acceptable answers by training on human preference data, which every large lab buys. Anthropic and Google run comparable human-review programmes.

Anonymised is not the same as unidentifiable

Prompts reach reviewers without account names attached, and OpenAI says it tries to strip personal information before they arrive. The company acknowledges that sensitive details can still come through.

The memory feature is a sharper problem. 404 Media reports that reviewers can see the memory summary blocks the product builds about a user, covering previous usage, location and life circumstances. A name is not required to identify someone described by where they live, what they do and what they have been asking about for months.

Michal Luria, a researcher quoted in the reporting, named the underlying mismatch: "Current chatbot interfaces automatically create a false sense of intimacy and privacy." People type things into a chat box that they would not put in an email, because the interface behaves like a confidant rather than a service with contractors attached.

HundredsContractors reading conversations
1 to 7Scale used to grade answers
$50+/hrReported pay for one worker
On by defaultTraining on chats for consumer plans

What OpenAI said, before and after

OpenAI has not hidden this exactly. After publication it pointed 404 Media to a help section stating that humans may review content to improve model performance, and updated that page with more detail about opting out.

The issue is the gap between a sentence in a help centre and the practice it describes. "Humans may review content to improve model performance" is true of a spot check by staff engineers and equally true of hundreds of contractors at an outsourcing firm reading conversations end to end and scoring them. A user reading that sentence would not picture hundreds of contractors reading conversations end to end. The contractor quoted by 404 Media agreed: asked whether users realise their chats are being reviewed, the answer was no.

The setting that decides whether this reaches you

On consumer plans, using conversations to improve models is on by default. With more than 900 million ChatGPT users, that default does most of the work.

Turning it off is a setting in the product rather than a support request, and the control governs whether conversations feed training and review, not whether they are stored. Business and enterprise arrangements generally sit outside consumer training defaults, which is the practical reason work conversations belong in a work account rather than a personal one.

Advertisement

The opt-out is also narrower than it sounds. It governs what happens to conversations from the point it is set, not what has already been collected, reviewed or folded into a model during training. Treating it as a delete button asks the setting to do something it does not do.

What to watch

Watch whether the disclosure changes in substance, not just in the help centre. A notice that says conversations may be read by external contractors, shown where people actually type, would be a different product decision from a paragraph added to a support article after a story ran.

The memory feature matters just as much. If reviewers see memory summaries, the anonymisation claim is weaker than it appears, and the fix is technical: strip the summary, not just the name.

Finally, watch the regulators. Data protection authorities in this region have been active on exactly this type of processing, and Indonesia, Malaysia and the Philippines have all moved this month on what platforms owe users about how their data is handled. A default-on training setting with external human review is the kind of arrangement those regimes were written to examine.

Advertisement
Priya Nair
Data, AI Governance & Policy Analyst

Priya Nair covers AI governance, data protection, privacy, and digital trust topics for RECATOOLS.

View author profile → · Editorial policy

About this byline Priya Nair is a RECATOOLS editorial persona for AI governance, privacy, and digital trust coverage. Articles are produced and reviewed under RECATOOLS editorial supervision.

Corrections policy

Advertisement