OpenAI is paying hundreds of contractors — one of them more than $50 an hour — to read real ChatGPT conversations and grade the chatbot's replies on a scale of 1 to 7, according to leaked internal documents obtained by 404 Media. The program has an internal codename, Project Lily, and a workforce recruited through a firm called Crossing Hurdles and paid out through the AI-training marketplace Mercor. What it does not appear to have is a disclosure that a typical ChatGPT user would ever have encountered. Asked by 404 Media to point to where it had told users that humans may read their chats, OpenAI did not answer. After publication, it pointed to a help page.

The work itself is unremarkable by the standards of the industry. Per the training materials 404 Media reviewed, a reviewer reads a user's prompt, writes a summary of what the person was trying to accomplish, then critiques and rates multiple candidate responses. The grading rubric is aimed at the model's personality problems: less sycophancy, less "AI-speak," fewer emoji, no anthropomorphizing itself. Reviewers are trained to flag lists stuffed with green checkmarks. This is reinforcement learning from human feedback working exactly as described in the literature — the part of model training that cannot be automated, running at the scale a chatbot with more than 900 million weekly users requires.

The privacy question is narrower and sharper than "humans read your chats," which has been true of consumer software since at least the Skype and Siri transcription stories of 2019. It is what those humans see. Contractors do not get usernames. OpenAI routes conversations through an automated tool it calls Privacy Filter, which replaces names, addresses, email addresses and phone numbers with labels. OpenAI's own documentation for that model concedes it can make mistakes, miss uncommon identifiers, and under-redact when context is thin — and 404 Media reports the company acknowledged sensitive details can still reach reviewers. Separately, The Next Web reported that the version of a chat handed to a reviewer can carry a "user memories summary" above the prompt: a digest of what the person has used ChatGPT for before, and sometimes where in the world they appear to live. A redaction filter that strips the name while preserving the interests, the history and the approximate location is doing less work than the word "anonymized" implies.

Then there is the disclosure. OpenAI's Data Usage for Consumer Services FAQ says that "a limited number of authorized OpenAI personnel, as well as trusted service providers that are subject to confidentiality and security obligations, may access user content only as needed," listing four purposes — abuse investigation, support, legal matters, and "to improve model performance (unless you have opted out)." That is a real disclosure, and it is accurate. It is also a help-center bullet point that describes a limited number of authorized personnel, which is a fair description of a security team and a strained one for hundreds of contractors reading a continuous stream. One of them put the gap plainly to 404 Media: "I don't think they would imagine some contractor somewhere [...] is analyzing the conversations."

The defaults do the rest. The "improve the model for everyone" setting is on by default for Free, Plus and Pro accounts, and off by default for Enterprise, Business and Edu. Turning it off applies to new conversations, not old ones. OpenAI spent August previewing zero data retention for enterprise customers. Consumers get the toggle that feeds the model, switched on, in a settings submenu.

Why it matters

Nothing here is unique to OpenAI, and the story is weaker if it pretends otherwise. Anthropic confirmed to 404 Media that it uses human review too — for users who switch the setting on, with account identifiers removed first. Google's Gemini carries a line saying humans review some saved chats. Human-in-the-loop review is how these products get less sycophantic and less robotic, and a version of it is disclosed in most major AI privacy policies.

What is specific to this story is the distance between the disclosure and the practice, and the fact that regulators have already ruled on that distance. The Court of Justice of the European Union held in September 2025, in EDPS v SRB, that a controller's duty to inform applies at the moment of collection and is judged from the controller's point of view, not the recipient's. Whether a contractor in North America could actually work out who wrote a prompt is therefore not the test. The obligation landed on OpenAI when it collected the conversation. Italy's Garante has already fined the company €15 million, €9 million of it for processing without an adequate legal basis, and ordered six months of public-information advertising on Italian television and radio.

The commercial context makes the exposure worse rather than better. The same product being fed to reviewers by default is the one OpenAI has spent the year pushing toward finance, health and other categories where users volunteer far more than they would type into a search box. A chatbot that behaves like a confidant generates exactly the training data that makes the next model better and exactly the training data that should not be sitting in a contractor's review queue.

What to watch

Whether OpenAI changes the disclosure rather than the practice — an in-product notice at the point of collection would resolve the EU question far more cheaply than restructuring the review pipeline. Whether any European regulator opens a formal inquiry citing EDPS v SRB, which would make Project Lily a test case rather than a news cycle. Whether OpenAI publishes numbers: how many conversations reach human reviewers, what share of accounts are sampled, and how often Privacy Filter fails. And whether the opt-out becomes retroactive, which is the single change that would most directly answer the complaint, and the one with the clearest cost to OpenAI's training pipeline.

“I don't think they would imagine some contractor somewhere [...] is analyzing the conversations.”
— A Project Lily prompt reviewer, OpenAI contractor, speaking to 404 Media
$50+/hr
Reported pay for one Project Lily contractor
1 to 7
Scale reviewers use to score ChatGPT replies
900M+
ChatGPT weekly users
EUR 15M
Italian Garante fine already levied against OpenAI