stackery.co

McDonald's AI Drive-Through: How 260 Chicken McNuggets Ended a Three-Year Experiment

McDonald's ended its three-year IBM voice AI partnership after order accuracy plateaued at 80-85%—worse than human workers. Over 100 locations reverted to headset operators by July 2024 after viral failures including a 260-nugget order.

Last updated 2026-09-05

The background

McDonald's runs drive-throughs at scale. Thousands of locations, millions of transactions, a business model built on speed and consistency. When a customer pulls up to the speaker, the interaction needs to be fast, accurate, and capable of handling the inevitable "actually, cancel that" moment that happens in roughly every third order. For decades, this meant a human with a headset, taking orders while managing three other tasks simultaneously.

The promise of voice AI in this context was obvious enough. Automate the order-taking, free up staff for other work, maintain consistency across shifts and locations. No sick days, no training variance, no accidental up-sell failures. The technology had been improving for years, and speech recognition had become genuinely good at transcribing what people said. On paper, a drive-through order is a bounded problem: limited menu, structured conversation, clear outcomes. Exactly the kind of task that AI should handle well.

McDonald's partnered with IBM to pilot a system called Automated Order Taker. The partnership made sense from both sides. IBM had been investing heavily in AI capabilities, and McDonald's had the scale to test whether voice automation could work in one of the most demanding customer-facing environments in retail. Over 100 US locations began testing the system. The pilot ran for roughly three years, long enough to work through initial problems and reach a stable level of performance.

This was not a hasty experiment. Three years is sufficient time to train models, refine prompts, handle edge cases, and demonstrate improvement. The scale was meaningful. The problem space was well-defined. And yet, by June 2024, McDonald's announced the end of the partnership. By 26 July 2024, the system was switched off everywhere it had been deployed, and over 100 locations reverted to human headset operators.

What actually happened

The Automated Order Taker went live at McDonald's drive-throughs starting in 2021. The system listened to customer orders, interpreted what it heard, and added items to the transaction. When it worked, it worked exactly as intended. When it failed, it failed in ways that were genuinely difficult to recover from.

The problems became visible not through internal metrics, but through customer documentation. People began recording their interactions with the AI and posting them to TikTok and Twitter. The videos spread widely, because the failures were not subtle. These were not minor errors in a receipt reviewed later. These were real-time breakdowns in a conversation, captured on video, often with the customer's incredulous commentary layered over the top.

In the most widely shared clip, the system kept adding Chicken McNuggets to a single order. Not mishearing the quantity. Not selecting the wrong item. Adding nuggets, repeatedly, while the customer tried to stop it. The order reached 260 nuggets before the video ended. Two hundred and sixty. The customer was filming, the system was still going, and there was no obvious way to make it stop. It was not a transcription error. It was a failure in the conversation model itself.

Another widely circulated failure involved bacon added to an ice cream order. Again, not a plausible mis-hearing. The system had lost the thread of what it was ordering and began adding items that made no sense in context. These were not edge cases in the statistical sense. They were the kinds of errors that humans simply do not make, because humans understand what an ice cream order is.

The pattern that emerged across many of the videos was the same: the system could not handle corrections. A customer would say something, the system would interpret it, and then the customer would attempt to clarify or change their mind. The system treated each new utterance as a fresh instruction rather than a modification of what had come before. "No wait, make that two" is a sentence that human order-takers process without conscious effort. For the AI, it was a new intent, disconnected from the previous one, leading to duplicate items or accumulating quantities.

Through 2024, the videos continued to circulate. The failures were becoming the public face of the experiment. Internally, McDonald's had data. Order accuracy had plateaued at roughly 80-85%. That figure had been stable for some time. It was not improving meaningfully, despite three years of operation and optimization. And crucially, it was worse than the humans it was meant to replace. Human drive-through workers typically achieve 90% or higher.

In June 2024, McDonald's announced the end of the IBM partnership. By 26 July, the Automated Order Taker was switched off in all test restaurants. Over 100 locations went back to the way they had been doing it before. The experiment was over.

The people in the room

No one involved in this project set out to build something worse than what already existed. The engineers working on voice recognition built a system that could transcribe speech accurately. The product teams defined success metrics. The operations people at McDonald's identified locations for the pilot and monitored performance. Everyone involved had reasonable expectations about what the technology could do, based on demos and controlled tests that likely showed much better results than the eventual field performance.

The decision to proceed with a multi-year, 100-plus location pilot was not reckless. Three years is enough time to demonstrate improvement. The problem space was genuinely suitable for automation. Speech recognition technology had advanced significantly. The business case was clear. And the initial results were probably encouraging enough to justify continuation. 80% accuracy might have seemed like a good starting point, with the assumption that it would improve toward parity with human performance.

What likely happened is that the comparison benchmark was not clearly established at the outset. When you are building something new, it is easy to measure it against nothing rather than against the thing it is replacing. Humans achieving 90% or higher order accuracy was a known figure, but perhaps not the figure that the project was being measured against. The system improved from its own baseline, plateaued, and by the time it became clear that it was not going to reach human performance, three years had passed and the failures had gone viral.

The damage

What actually went wrong

Voice ordering in a drive-through is hard, but not for the reason most people assume. The speech recognition component—the part that converts sound into text—is actually quite good. Modern AI can transcribe speech accurately even in noisy environments. The Automated Order Taker could hear what customers said. That was not the problem.

The problem was the correction loop. Human conversation, especially transactional conversation like ordering food, is full of modifications, clarifications, and changes of mind. "Can I get a Big Mac. Actually, make that two. No wait, just one, and add a large fries instead." A human order-taker parses this without effort. They understand that each utterance modifies the previous state of the order. They maintain a mental model of what has been said and what has been changed.

The AI system was built around intent classification. Each time the customer spoke, the system tried to determine what they wanted and translate that into an action. But it had no reliable model of what the customer had already changed their mind about. "Make that two" is not an intent on its own. It is a reference to something previously discussed. The system treated it as a new instruction, leading to duplication. "No wait" should signal cancellation of the previous action, but the system heard it as a new utterance and tried to find an intent in it.

This is why the nugget video is important. Two hundred and sixty nuggets accumulating in real time is not a speech recognition failure. The system was hearing the customer. It was interpreting what it heard as instructions to add more nuggets, because it had no mechanism to understand "stop, that is too many, go back." That is a fundamental architectural problem, not a training data problem.

The accuracy figure is worth examining closely. Roughly 80-85% sounds respectable in isolation. Four out of five orders correct. But it was worse than the humans it was replacing, who managed 90% or better. A system that performs worse than the incumbent is not an early version of a better product. It is a worse product. And the failure distribution mattered. The errors were not evenly spread across minor mistakes. Some of them were spectacular, film-worthy breakdowns that shaped public perception far beyond their frequency.

When your automation fails in front of customers, your customers become your QA team. And they publish their findings. Not in a bug tracker, but on TikTok, where millions of people watch a drive-through order spiral into absurdity. That is a different kind of damage than a line item in a performance dashboard.

What small businesses can learn

Sources

Get the shortlist, not the noise

One email a week. The tool we would actually buy, and why.

Join the newsletter