Amazon's Math: Cut 30,000, Hire 10,000
Employees gaming an internal AI usage leaderboard exposed the cracks in that logic.

Opening
Hello, reader. This is OZ Talking. Last week, AWS CEO Matt Garman said something in an interview that stuck with me. “Replacing entry-level employees with AI is the dumbest thing I’ve ever heard.” In the same interview, he revealed that Amazon is hiring 11,000 interns and entry-level employees this year. I was floored when I heard it.
Just months before he said this, Amazon had cut about 30,000 corporate jobs. Amazon CEO Andy Jassy sent employees an official memo stating that AI would reduce the company’s overall corporate workforce. At the same time, Amazon is pursuing plans to replace more than 500,000 jobs with robots.
It’s a strange calculus—hiring and cutting jobs at the same time. But what caught my attention wasn’t this contradiction itself. It was something else happening inside Amazon: employees deliberately gaming an AI usage leaderboard, in what’s come to be called “tokenmaxxing”1.
🏷️ Faking It: The People Who Pretend to Use AI
Here’s what the Financial Times reported this past May. Inside Amazon, there was a leaderboard called KiroRank. It was a dashboard that ranked employees by how many tokens2 they consumed using Kiro, Amazon’s internal AI coding tool. Amazon had set a goal for more than 80% of its developers to use the AI tool at least once a week.
The result was predictable. Employees used an internal agent3 tool called MeshClaw to automate tasks like deploying code, sorting emails, and processing Slack messages—except many of these tasks didn’t actually need doing. They were feeding the AI busywork just to climb the rankings. Employees called this “tokenmaxxing.”
One employee told the FT, “There’s enormous pressure to use this tool. Some people are maximizing their token usage with MeshClaw.” Amazon said token usage wasn’t factored into performance reviews, but employees felt that managers were unofficially checking the data anyway. Security concerns surfaced too. MeshClaw had permission to deploy code and interact with internal systems on a user’s behalf. One employee said the default security settings were “terrifying” and that they “couldn’t just let it wander around on its own.” I’ve actually covered a similar case at Uber in this newsletter before.
Told to Spend Half Your Salary on Tokens… How’d That Go?What happened after Uber burned a year’s budget in four monthsThis wasn’t just an Amazon problem. At Meta, a leaderboard called “Claudeonomics” tracked token consumption across 85,000 employees, and over 30 days, they burned through 60 trillion tokens—worth roughly $9 billion at public API pricing. Top users were awarded titles like “Token Legend” and “Session Immortal.” Some employees left AI agents running research tasks for hours, unattended, just to climb the rankings. The leaderboard was shut down within 48 hours of the story breaking.
At Microsoft, President Julia Liuson sent an internal memo declaring that “AI usage is no longer optional—it’s core to every role, at every level.” Nvidia’s Jensen Huang went even further, publicly saying he’d be “deeply concerned” if a $500,000-salary engineer wasn’t burning through $250,000 worth of tokens.
Here’s the interesting part: all three companies shut down or restricted access to their leaderboards at almost the same time. Across Big Tech, AI usage had become a performance metric—and then, almost simultaneously, everyone quietly admitted that metric had failed.
📊 Is the Demand Behind a $700 Billion Bet Real?
Garman summed up the episode this way in the interview: “You have to measure what you actually want to measure. We built the wrong metric, and people found ways to game the metric instead of pursuing the goal.”
He’s right. But a more fundamental question remains: why was this metric created in the first place?
This year, the combined AI capital expenditure (CAPEX)4 of Amazon, Alphabet, Meta, and Microsoft is roughly $700 billion—up more than 70% from $410 billion last year, or about ₩1,000 trillion in Korean won. The core premise behind this spending is that AI demand will keep outstripping supply.

But what if some of that demand was just tokenmaxxing? Of the 60 trillion tokens that 85,000 employees burned through in 30 days, how much was real work, and how much was leaderboard gaming?
External, paying-customer demand is a separate matter, of course. In fact, in a CIO survey Garman cited, about 90 out of 100 respondents said they were seeing ROI on their AI investments. He pointed to cases where insurance claims processing times dropped sharply from 60–90 days, and where agents pushed workflow success rates from 20% to over 90%.
And Amazon is genuinely shipping AI products that work. Amazon Connect Talent, launched this past April, is a hiring tool where AI conducts voice interviews autonomously, around the clock. Garman said, “We want our recruiters focused on finding candidates directly and building relationships,” adding, “entering details into a form isn’t the work they want to be doing.” So here’s a company selling a tool that lets AI conduct job interviews, while its executive says replacing entry-level hires with AI is dumb. The two statements look contradictory, but inside Garman’s logic, they coexist just fine: if AI handles the repetitive work, humans move up to higher-value work.
But Garman himself conceded something: “Just as many companies—if not more—are killing PoCs5 that showed no results.” The point isn’t that all this demand is fake. It’s that inside the usage data underpinning hundreds of billions of dollars in infrastructure spending, performance signaling and actual value creation are mixed together, indistinguishable.
There’s a concept in economics called Goodhart’s Law6. The principle is: once a measure becomes a target, it stops being a good measure. The moment token consumption was set as a proxy for AI productivity, it lost its ability to actually measure AI productivity. Tokenmaxxing is a textbook case of this law in action.
🔍 The Ladder Is Being Pulled Away in Korea, Too
Let’s pause on Garman’s “Excel analogy.” He said, “Excel eliminated jobs based on manual calculation, but people learned computers, and the labor market expanded.” He’s not wrong. But this time, the structure is a bit different.
A Stanford research team’s analysis of ADP payroll data found that in occupations with high AI exposure, employment for early-career workers aged 22–25 fell by about 20%, while employment for workers in their 30s and older actually rose by 6–9%. AI isn’t eliminating all jobs—it’s automating the specific work junior employees use to build experience: summarizing, organizing, drafting. In other words, the bottom rungs of the ladder.
Korean data shows a similar pattern. According to Catch, a Korean job-posting platform, entry-level IT and telecom job postings at large corporations fell 67% year-over-year. Overall entry-level full-time hiring at large corporations dropped 43%. This past May, the number of regular employees in Korea turned negative year-over-year for the first time in 26 years and 5 months. In the information and communications sector alone, regular positions held by workers in their 20s fell by 57,000, while the same category for workers in their 30s grew by 26,000. Hiring structures are shifting away from entry-level and toward experienced hires.
When Excel arrived, calculators could “climb” into accounting roles. But what AI is automating now is the rung itself—the process of climbing. It’s the boilerplate code junior developers write while learning how systems work; it’s the data cleanup junior analysts do while absorbing business context. In the words of the Stanford Social Innovation Review, what AI has consumed is “the low-risk, repetitive work that used to teach people how to operate inside an organization.” When this “learning work” disappears, a structural gap opens in the talent pipeline, and the pool of people ready to become mid-level managers and senior staff in three to five years starts drying up.
Some companies have recognized the problem. IBM concluded that relying solely on AI efficiency isn’t sustainable long-term and announced it would triple entry-level hiring. IBM’s CHRO said, “The companies that succeed most in three to five years will be the ones that doubled entry-level hiring right now.” Cognizant hired 20,000 entry-level employees in 2025 alone.
Garman’s comment that “the most important skill going forward isn’t any particular technical skill but a willingness to learn” needs to be read in this context. He’s right. But in a structure where opportunities to learn are themselves shrinking, “willingness to learn” alone can’t substitute for the ladder.
Oz’s Lens
In my work building GTM and technology strategy for companies, I’ve confirmed the same thing over and over: measurement shapes behavior. Set MAU as your KPI, and employees focus on sign-ups instead of retention. Measure commit counts, and meaningless commits pile up. Measure token consumption, and tokenmaxxing is the inevitable result.
What interests me more is how this phenomenon dovetails with the structural incentives of infrastructure providers. AWS sells AI infrastructure. It only sells servers if customers use a lot of AI. So when Garman says jobs change rather than disappear, he might genuinely believe it—but he’s also in a position where he has to say it. If customers get scared of adopting AI, they stop buying infrastructure.
This isn’t a problem specific to Garman. It’s a structural double bind that every infrastructure provider inevitably carries: they have to say “AI will change the world,” while also saying “but the world shouldn’t change too much.”
The real test of AI adoption is simple. The question should be not “how much are we using it,” but “what actually changed”. If insurance claims processing dropped from 60 days to 3, that has value regardless of how many tokens it took. This is exactly why, as Korean companies push forward with AI transformation, they need to design outcome metrics before usage metrics.
Closing
Here’s the summary.
Amazon cutting 30,000 jobs while hiring 10,000 isn’t a contradiction—it’s a kind of portfolio swap. But whether that swap is grounded in real productivity, or standing on the same inflated “usage” numbers as tokenmaxxing, is still an open question. As Big Tech pours $700 billion into AI this year, and Korean companies follow with their own AI transformation spending, the simplest question for testing whether that demand is real is this: not “how much AI are we using,” but “what changed because of it.”
Have you ever been told at work to track “usage rate” or “adoption” after rolling out an AI tool? Looking back, were you actually measuring the tool’s value—or just evidence that it was being used? Share your experience in the comments, and I’ll bring it into a future issue.
💬 Share your thoughts or experience on measuring AI adoption in the comments · 📨 If this piece might help a colleague, please share it.
References & Further Reading
Primary sources
- Casey Newton, “The CEO of AWS on why Amazon is hiring 11,000 interns and junior employees”, Platformer, 2026.6.24. : Contains Garman’s direct remarks on the Excel analogy and the token leaderboard.
- Financial Times, “Amazon employees inflate AI usage on internal leaderboards”, 2026.5.11. : The original report on the tokenmaxxing phenomenon, including interviews with Amazon employees.
- The Information, “Meta’s Claudeonomics leaderboard”, 2026.4.7. ··· The original report on Meta’s 60-trillion-token consumption.
Background
- Erik Brynjolfsson, Danielle Li, Raymond Wang, “Canaries in the Coal Mine? Six Facts about the Recent Decline in Entry-Level Hiring”, Stanford Digital Economy Lab, 2025. : An empirical study analyzing the decline in employment among early-career workers aged 22–25. This is the core evidence behind the “bottom rungs of the ladder” discussion in this issue.
- Stanford HAI, “AI Index Report 2026”, 2026.4. : An annual report compiling key data, including the 20% decline in junior developer hiring and a 53% AI adoption rate.
- Catch (a Korean job-posting platform), “2025 Large Corporation Entry-Level Hiring Postings Analysis”, 2025.12. : The source of the data on the 67% plunge in entry-level IT/telecom hiring postings at large Korean corporations.

The author, Kwangseob Ahn, is a professor of business administration at Sejong University and lead consultant at OBF (Oswarld Boutique Consulting Firm). He teaches statistics and data analysis — business data management and business analytics — while leading GTM and AI strategy consulting in the field, designing the seam between technology and business. He has published academic research on a memory architecture for AI dialogue systems (HEMA) and runs Daily Arxiv, a daily curation of global AI papers. He holds a master’s from Korea University’s Graduate School of Technology Management and a KMBA. He is the author of Homo Brainless: The People Who Outsource Their Thinking.
Footnotes
-
Tokenmaxxing: The practice of artificially inflating AI tool usage to climb internal leaderboard rankings. A portmanteau of “token” (the unit of data AI processes) and “maxxing” (maximizing), this phenomenon emerged simultaneously across Big Tech in 2026. ↩
-
Token: The smallest unit of data an AI model uses to process text. One Korean character typically corresponds to roughly 2–3 tokens, and most AI service pricing is based on token consumption. ↩
-
Agent: AI software that can make decisions and carry out tasks on its own, without human intervention. Unlike a simple chatbot, an agent can autonomously handle practical work like deploying code, processing emails, or managing schedules. ↩
-
CAPEX (Capital Expenditure): Spending a company makes on long-term assets like data centers, servers, and real estate. In the AI era, building GPU servers and power infrastructure accounts for most of CAPEX. ↩
-
PoC (Proof of Concept): A small-scale experiment to test whether a new technology or idea actually works. If it succeeds, it moves toward full-scale deployment; if it fails, it’s discontinued. ↩
-
Goodhart’s Law: A principle proposed by British economist Charles Goodhart, stating that once a measure becomes a target, it stops being a good measure. A prime example: the moment token consumption shifted from being an “indicator” of AI productivity to being the “target” itself, it lost its ability to measure productivity at all. ↩

Your take shapes the next issue
What resonated most in this issue, or where has your experience been different?