SEATTLE, 27 AUG 2026 — Amazon is closing Mechanical Turk on 30 September, twenty-one years after launching the marketplace that let companies hire people, a few cents at a time, to do the work software could not. SageMaker Ground Truth, its enterprise data-labelling product, goes with it.
The headline writes itself and it is wrong. Artificial intelligence did not make the humans redundant. It made them more expensive.
What is closing
Mechanical Turk opened in 2005 as a marketplace for what Amazon called Human Intelligence Tasks. These were the countless small judgements a program cannot make, from transcription and image labelling to survey completion and content moderation. At its peak it drew on more than 500,000 workers.
Amazon says the decision followed an internal assessment of its programmes and services. It stopped accepting new customers last month, which workers read at the time as the signal it turned out to be.
Named after a fraud, and that was the point
The original Mechanical Turk was an eighteenth-century chess-playing automaton that toured Europe defeating opponents, and it worked because a human chess master was concealed in the cabinet.
Amazon knew exactly what it was invoking. Jeff Bezos described the service as artificial artificial intelligence. The joke was precise. A program could call a human through an interface indistinguishable from calling a function, and the calling code neither knew nor cared that a person was on the other end.
That idea did not die with the platform. It is the direct ancestor of every reinforcement learning pipeline that ranks model outputs by human preference, and of every evaluation harness where people score answers a model has produced. The cabinet got bigger and the interface got better. Someone is still inside.
The obvious reading is wrong
The easy conclusion is that models became good enough to do the work themselves. The evidence points somewhere else.
Demand for human judgement in AI development has grown, not shrunk. What changed is which judgement is worth buying. When the task was drawing a box around a car in a photograph, the cheapest competent worker anywhere in the world was the right supplier, and a marketplace of anonymous piecework was the efficient way to reach them.
The tasks that matter now are different. They require expertise: deciding whether a model's answer to a clinical question is safe, whether a legal summary misstates a holding, whether generated code carries a subtle defect. You cannot buy that in three-cent increments from a queue.
So the market did not disappear. It stratified. Amazon built for the bottom of it, competitors built for the top, and the bottom is the part a model can now do.
What the market moved to
Scale AI, Mercor and Prolific are the names that drew the work away, and the shift in what they sell describes the change better than any statement from Amazon.
Where Mechanical Turk offered volume at the lowest available price, its successors compete on the credentials of the people they recruit — physicians, lawyers, engineers, doctoral researchers — and pay accordingly. That is a different business with different economics, and Amazon appears to have decided it was not one worth entering late.
The old model had its critics. Effective hourly earnings on task marketplaces were persistently documented as low, with unpaid time spent searching for work and no recourse when a requester rejected a completed task. The platform's closure is not a straightforward loss for its workers, though it is a loss of income for some. Both things are true, and most coverage picks one.
The constituency nobody is writing about
Less discussed is Mechanical Turk's role over two decades as a backbone of academic research.
Psychology, behavioural economics and survey methodology leaned on it heavily, because it solved a problem that had constrained those fields for a century: recruiting several hundred participants in an afternoon, at a cost a departmental budget could absorb, without a campus subject pool of undergraduates. An enormous body of published work rests on samples gathered there.
That came with well-aired methodological arguments about how representative such samples are, about participants who take so many studies that they recognise the manipulations, and about attention checks that filter for compliance as much as for care. Those debates ran for years and shaped how the field reports its methods.
Prolific, one of the services named as drawing workers away, grew specifically out of that dissatisfaction, offering researchers better-characterised participants at higher cost. So the academic migration was already underway. What the closure does is end the cheap option, and departments in places where research funding is thin will feel that more sharply than a laboratory in Boston.
Ground Truth is the larger signal
Closing the public marketplace might look like tidying up a declining consumer product. Closing SageMaker Ground Truth, the enterprise offering, signals something larger.
Ground Truth was the enterprise offering, the managed way for an AWS customer to get training data labelled inside the platform where the training happens. Retiring both means Amazon is leaving the data-labelling layer, not merely retiring an old brand.
That is a strategic judgement about where the value sits. Amazon is investing enormously in the compute underneath models and in the applications above them, and has concluded that the layer in between is someone else's business. Given that this is the layer everyone agrees is the bottleneck on model quality, it is a striking place to decline to compete.
What it means from here
For workers in this region the immediate effect is concrete. Task marketplaces have long drawn substantial participation from South and Southeast Asia, where the pay compared favourably with local alternatives even where it compared badly with a minimum wage, and that particular door closes on 30 September.
The work available in its place asks for verifiable expertise. That is better paid and reaches fewer people, and it favours those who can demonstrate a credential to a platform's satisfaction — which is a meaningfully different population from the one that could complete tasks on a phone between shifts.
The broader point applies to any AI system described as autonomous. A support agent resolving contacts with no human in the loop was trained and evaluated by people, and continues to be. The humans did not leave the loop. They moved further up it, out of view, and got harder to count.