BusinessIssue #108

Why Starbucks Pulled AI Out of Its Stores

It promised 99% accuracy, but baristas trusted their own hands more.

Why Starbucks Pulled AI Out of Its Stores

Opening

Dear reader, picture a bottle of peppermint syrup sitting on a shelf. The vanilla syrup right next to it gets recognized perfectly, but this one bottle is processed as if it doesn’t exist. That’s the real face of the AI inventory system Starbucks in the US used to boast about.

On May 19th, Starbucks completely scrapped the AI inventory-counting tool called Deep Brew, which it had rolled out to 11,000 stores across North America. Just 9 months after adoption. To cut to the conclusion: this isn’t a failure of AI technology — it’s a failure to ask whether AI was even the right tool for this problem in the first place.

Deep Brew: The Secret AI Engine Behind Starbucks’ 30% ROI — And Wha…After a brilliant conversation with Jemi Crookes I’ve been diving deep into how world-class brand…linkedin.com

Last year, Starbucks made a huge publicity push around bringing AI into its stores. In some sense, just a bit over half a year later, they’ve pulled the whole thing back out.

📦 What Happened at 11,000 Stores

In September 2025, Starbucks rolled out inventory AI from Seattle-based startup NomadGo to every store in North America. Using LiDAR1​ sensors and tablet cameras, the system was designed to automatically count syrups, milk, and beverage ingredients on the shelf. NomadGo claimed the system was 8 times faster than manual counting with 99% accuracy, and Starbucks CTO Deb Hall Lefevre introduced it by saying it would “let partners focus on crafting drinks and connecting with customers instead of counting inventory.”

But reality on the ground was different.

According to Reuters’ reporting, the system repeatedly confused similar-looking milk varieties and outright failed to recognize products that were clearly sitting on the shelf. Even in the promotional video Starbucks released at launch, a bottle of peppermint syrup went unrecognized between two neighboring products — caught on camera.

The biggest problem was that the trust threshold2​ collapsed. Employees had to double-check the AI’s output every single time. A tool introduced to replace manual labor ended up creating double the work instead. Previously, counting by hand once was enough. Now, AI counted first, and then a person had to count again. The tool didn’t reduce work — it just added cognitive load.

Let me add some context here. This tool was a core pillar of CEO Brian Niccol’s “Back to Starbucks” turnaround strategy. Niccol, who came over from Chipotle in September 2024, diagnosed inventory shortages as a major drag on sales and rapidly pushed this technology — which had been in testing since his predecessor’s tenure — into every store. In the meantime, North America’s operating margin had fallen from 18% two years earlier to 9.9%, and the turnaround needed speed.

NomadGo stated that over the course of 2025, the system counted more than 186 million items across 11,000 stores. The number sounds impressive on its own, but nobody has disclosed how many of those counts were actually re-verified by employees.

On May 19th, Starbucks announced the tool’s official discontinuation via a company-wide memo. “Effective today, automated counting is being discontinued. Beverage ingredients and milk will now be counted the same way as all other inventory items.” The blog post announcing the launch has already been deleted. Starbucks told Reuters it was “a decision to focus on consistency and execution across stores,” but the company never used the word failure.

🔍 It Wasn’t the Technology — It Was the Wrong Question

Why did it fail? Much of the coverage blames NomadGo, the AI company — framing this as a problem of “insufficient AI accuracy.” But in my view, there’s a more fundamental issue.

First, full-scale deployment without verification. NomadGo’s “99% accuracy” was a self-reported figure. It was rolled out across all 11,000 stores at once, with no independent third-party validation. Tech Times’ analysis nails the core issue: “The 99% accuracy claim was never independently verified before deployment across 11,000 stores.” What needs scrutiny here is the number 99% itself. If there are 20 items on a shelf, 99% accuracy means getting 0.2 of them wrong. Sounds fine, right? But multiply that across 11,000 stores, counted multiple times a day, and that 0.2 stacks up into tens of thousands of errors daily. And the real-world error rate fell well short of 99% anyway.

According to a RAND Corporation analysis of more than 2,400 enterprise AI projects, 80% of AI projects fail to achieve their intended business value. Of those, 34% are halted before reaching production, 28% are completed but fail to deliver the expected value, and 18% deliver some value but don’t justify the investment cost. MIT’s 2025 study is even more blunt: 95% of enterprise generative AI pilots produced no measurable revenue impact.

Gartner’s 2025 report likewise projected that 60% of AI projects lacking AI-ready data would be abandoned by 2026. According to S&P Global, large enterprises (10,000+ employees) abandoned an average of 2.3 AI projects in 2025, with an average sunk cost of $7.2 million (~₩10 billion) per abandoned project.

The breakdown of failure causes is also worth noting. In an analysis of 140 enterprise AI implementations, only 23% failed due to model performance or technical integration issues. The remaining 77% failed due to strategy, governance3​, and change management. It’s not a technology problem — it’s a people and organization problem.

Second, the problem itself was misdefined. Starbucks’ inventory problem was never caused by inaccurate “counting.” According to Reuters’ in-depth reporting earlier this year, less than a third of Starbucks deliveries arrived on time, and more than 1,500 cup-and-lid combinations were adding complexity to the supply chain. It wasn’t that manual counting lacked accuracy — the supply chain structure itself was the problem.

This is a pattern that repeated under the previous CEO too. Under CEO Laxman Narasimhan, Starbucks partnered with o9 Solutions to introduce an “automated ordering” system — a machine learning system that consistently recommended quantities lower than what was actually needed. The technology changed, but the same mistake — adopting a solution before defining the problem — stayed exactly the same.

🎯 How to Pick the Right Tool for the Problem

What the Starbucks case shows is a simple but easily forgotten principle: not every problem needs technology of the highest caliber.

A café shelf is not a controlled environment. Lighting shifts, product placement changes constantly, and similar-looking milk containers sit side by side. When a computer vision4​ system trained under uniform conditions gets deployed into a highly variable real-world setting, its performance can degrade sharply. A warehouse, where SKUs are fixed and shelf positions are standardized, is fundamentally different from a café shelf that baristas rearrange all the time.

Human hands, on the other hand, adapt flexibly to environments like this. A barista can instantly tell oat milk from low-fat milk at a glance, and immediately account for a shelf rearrangement from just yesterday. While an algorithm’s retraining5​ costs time and money, a person processes context in real time. What’s more, a person can simultaneously make the judgment “this syrup is almost out” — not just counting, but judgment that incorporates context.

Feedback from a Starbucks store employee captures this exactly: “Thank you for getting rid of automated counting. The intent was good, but it was hard to execute.”

The point here isn’t that “humans are better than AI.” It’s that choosing the right tool for the nature of the problem comes first. Let’s revisit Starbucks’ inventory-counting problem. For the task of distinguishing visually similar, low-volume products in a high-variability environment, there were several options on the table. Besides expensive computer vision AI, they could have installed weight sensors on shelves to detect changes in stock levels. A barcode scanner linked to a simple database might have sufficed. Or, as it turns out now, having a person count by hand might simply be the most accurate option.

On the other hand, for analyzing sales data across thousands of stores to optimize ordering patterns — that’s clearly a domain where AI does better. Even within the same problem of “inventory management,” the optimal tool differs at every stage.

To someone holding a hammer, everything looks like a nail. Once you start from the premise that “we must adopt AI,” it becomes easy to see every problem as something AI can solve.

Oz’s Lens

Honestly, the moment I saw this news, I thought of something I say constantly in my consulting work.

I currently consult for various companies on AI adoption, and I never unconditionally recommend deep learning or generative AI. In some cases, attaching a single sensor is far more efficient. There’s a surprising number of tasks that can be handled perfectly well with Excel VBA. And even when AI is the right call, a top-spec frontier model6​ is usually unnecessary. From embedding models7​ to lightweight classification models, deploying the right model in the right place — that’s what real expertise looks like.

There’s a pattern I’ve seen over and over while building go-to-market strategies. The binary framing that “digital is always superior to analog, agile always beats waterfall, horizontal organizations are always better than vertical ones.” Reality doesn’t work that way. Depending on the situation and the specific problem, there’s a tool and methodology that’s genuinely more efficient and better suited. Finding the most optimal fit — that’s the real task in front of us. Starbucks just proved this simple principle at the scale of 11,000 stores.

Closing

There are three things worth remembering from Starbucks scrapping its AI inventory tool.

One: deploying a vendor’s 99% claim at full scale without verification comes back to cost you double. Two: 77% of AI project failures stem not from technology but from strategy and problem definition. Three: there’s no need to use computer vision where a sensor will do, and no reason to bolt a frontier model onto a task that VBA already handles.

When your organization is considering adopting AI, ask this question first: “What technology does this problem actually need?” If you’re wrestling with that question, reach out to me. I’ll think through the best answer with you and deliver the optimal solution. Oswarld Boutique Consulting Firm is always open. :) contact@oswarld.com

References & Further Reading

Primary sources

Background

The author, Kwangseob Ahn, is a professor of business administration at Sejong University and lead consultant at OBF (Oswarld Boutique Consulting Firm). He teaches statistics and data analysis — business data management and business analytics — while leading GTM and AI strategy consulting in the field, designing the seam between technology and business. He has published academic research on a memory architecture for AI dialogue systems (HEMA) and runs Daily Arxiv, a daily curation of global AI papers. He holds a master’s from Korea University’s Graduate School of Technology Management and a KMBA. He is the author of Homo Brainless: The People Who Outsource Their Thinking.

Footnotes

  1. LiDAR (Light Detection and Ranging): a technology that measures the distance and shape of objects by firing lasers and reading the reflected light. It’s the same kind of sensor found on the back of iPhone Pro models for 3D spatial recognition.

  2. Trust Threshold: the minimum accuracy level at which a user can accept a system’s output without a separate check. Once accuracy falls below this line, humans must re-verify the system’s results, and the tool’s efficiency disappears.

  3. Governance: the decision-making system and management structure an organization uses when adopting or operating a technology or project. It covers who decides, by what criteria things are evaluated, and how problems are handled when they arise.

  4. Computer Vision: technology in which AI analyzes images or video captured by a camera to recognize and classify objects. Used in things like facial recognition and self-driving cars’ surroundings detection.

  5. Retraining: the process of further training an AI model so it can adapt to a new environment or new data. It’s required every time field conditions change, and it costs time and money.

  6. Frontier Model: top-tier, large-scale AI models like GPT-5.5, Claude, and Gemini. They perform extremely well, but aren’t suited to every task in terms of cost and processing speed.

  7. Embedding Model: a lightweight model specialized in converting text or images into numerical vectors to calculate similarity. It can be used faster and more cheaply than a frontier model for tasks like search, classification, and recommendation.