Why Big Tech Wants a Camera on Your Face
We asked for Samantha, and what showed up is an eye watching you.
Opening
There’s a face that comes to mind whenever we imagine artificial intelligence. Iron Man’s Jarvis greeting you with “Good morning,” or Samantha from the film Her picking up on your mood before you say a word. AI gets cast as something that understands you — a companion standing on your side, helping you through the day.
In July 2026, that fantasy is finally arriving as machinery. OpenAI’s first device, whose outline emerged in reports just days ago, is a screenless, speaker-shaped companion. It comes fitted with a camera and sensors, moves on its own just enough to feel “alive,” and even reads your email to get to know you better over time. Meta, meanwhile, is already selling cameras you wear on your face — smart glasses. Samantha and Jarvis have landed on the shelf.
But I read this news with a familiar unease rather than excitement. A few months ago, I covered a wealthy Toronto neighborhood wrapping itself in AI surveillance cameras, and I asked: the moment you hand Jarvis the keys, whose Jarvis does it become? Back then, the surveillance lived in the neighborhood, mounted on street poles. This time, that eye is moving onto your face, into your living room. And this time, we’re the ones inviting it in — because we want a companion.
This piece is about the fine print on that invitation — the hidden trade buried inside the wish for a Jarvis, and why Big Tech wants so badly to put a camera on your face.

Let me flag one thing up front. The AI industry is currently split. On one side, Mira Murati is betting on bigger, more tunable language models (the 975B-parameter open-weight model she released recently). On the other, Yann LeCun has left his company altogether to dig a different path, arguing that language models simply cannot understand the world. The directions are opposite, but both roads arrive at the same destination — something that has to watch you.
What’s Being Built Isn’t a ‘Smart Assistant’ — It’s an ‘Always-On Observer’
OpenAI’s first hardware device is reportedly a screenless speaker that can move | TechCrunchThe device is weirdly described as involving “mechanical elements that can move on their own” and the Bloomberg report includes the detail that the device is designed to “feel like a companion and becBreak down the specs of OpenAI’s first consumer device, and you’ll find almost no talk of performance. Instead, you get this: a camera and ambient sensors, a battery you can carry from room to room, mechanical parts that move on their own to fake being alive, and a personality that grows more personalized and more proactive the longer it watches you. The battleground the company is staking out isn’t benchmark scores — it’s “personality” and “presence.” That’s also why it paid $6.5 billion for a hardware team called io and borrowed the hands of Jony Ive, the man who built the iPhone: to make the technology feel personal.
Here’s an uncomfortable symmetry: the specs for a companion that looks out for you and the specs for a surveillance device that watches you constantly are effectively identical. Anticipating your needs requires observing your daily life; reading your email is what makes something “know” you well. Convenience and surveillance are built from the same materials.
Meta Ray-Ban Display – The Innovative Future of AI Glasses - haebomMeta has fired the starting gun on a new era of next-generation AI glasses — built on the Ray-Ban Display and Neural Band with EMG technology — poised to replace the smartphone.Meta has been pushing in this direction much earlier and much harder. Screenless smart glasses grew 167% by unit volume in Q1 2026 alone, making them the fastest-growing category in wearables, with Meta and EssilorLuxottica together controlling more than 80% of the market. Last month, rather than hiding behind the Ray-Ban brand, Meta put its own name on a pair of glasses priced at $299 — $80 cheaper than its existing line. The strategy: seed volume first, then widen the moat with the data that follows. The headline feature the glasses tout is “super sensing” — the ability to recognize objects and places in front of you in real time.
To sum up: the competition has shifted from “who’s smarter” to “who’s always closer, seeing more.” That the winning form factor happens to be a camera on your face is no accident.
For Jarvis to Be Useful, It Has to See the World — Which Means Seeing You
Let’s go one layer deeper. Why is money pouring into camera-equipped hardware right now, of all times? The answer lies on the model side.
For Jarvis to be truly useful, being articulate isn’t enough. It has to understand the layout of your room, where objects sit, what’s happening right now. Yann LeCun believes today’s AI simply can’t do this. He recently stepped down as Meta’s chief AI scientist to found a company called AMI Labs (roughly €500 million in scale), and his founding thesis was exactly this point: today’s models only look smart because they handle language plausibly — they don’t actually understand the world.
The alternative he’s pushing is the world model1 and JEPA2. The core idea: the laws of physics aren’t in text. No amount of reading the internet will tell you how a cup tips over or how light falls on a surface. That knowledge lives in the world. So this camp learns from video instead of text — by “watching” the world. Fei-Fei Li, who has raised $1 billion leading the world-model startup World Labs, put it this way: language models are “linguistic artisans working in the dark” — skilled with words but blind to the world.
This is the crux of the piece. Research aimed at making AI “understand the world” is, mechanically speaking, research aimed at “observing the world constantly.” Modeling the world requires watching the world, and you happen to be inside that world. The perception layer and the surveillance layer are, in fact, the same layer.
The fact that LeCun bothered to leave his company reads as a signal in itself: the world-understanding capability that would serve as the companion’s “brain” isn’t ready yet. Because the brain is lacking, more observational data is needed; because more observational data is needed, sensors have to go on your face and in your living room. The hardware race and the world-model race aren’t two separate stories — they’re two halves of the same machine.
But This Isn’t Happening Because Consumers Want It
Let me clear up a common misconception here — the narrative that “chatbots are over, everything’s moving off-screen now.” I think that’s overstated. The evidence, in fact, points the other way.
The immersive push to put people inside virtual worlds is collapsing. Apple has scrapped development of an affordable XR headset display it was weighing as a Vision Pro follow-up (an effort known in the industry as G-VR), and Samsung Display has decided to wind the project down early, in September. The $3,499 Vision Pro has faced weak sales alongside reports of discontinuation and halted production. Consumers didn’t open their wallets for the promise of a screen filling their field of view and transporting them to another world.
Apple Scraps ‘Budget XR Display’ Development - THE ELECApple is halting development of the affordable extended reality (XR) device panel project it had been reviewing as a follow-up to the Vision Pro. As the hardware trend shifts from XR headsets to AI smart glasses, the harWhat matters is the scene that follows. Big Tech hasn’t abandoned wearables — it has swapped “immersion” for “observation.” Instead of a headset that puts you inside a virtual world, it’s building glasses that bring AI into your reality. Why the pivot? Not because demand is red-hot. Demand for an always-on companion device has never actually been proven. There’s already a graveyard of products that sold a similar dream and vanished — Humane and Rabbit among them.
Yet there’s one reason they keep pushing: whoever controls the interface after the smartphone controls the next platform. Meta keeps holding onto glasses even while absorbing roughly $10 billion a year in losses at Reality Labs. OpenAI needs a hardware narrative to show investors ahead of an IPO. That $299 price tag isn’t aimed at profit — it’s aimed at market share, laying the groundwork before Google and Samsung bring Android XR glasses to market this fall.
So the subject of what’s happening right now isn’t “you” — it’s Big Tech. The market isn’t moving because you want Samantha; Big Tech wants the throne of the next platform, and it’s selling you the story of “Samantha” to get there.
The Virtual World Has Already Rehearsed This
If this observation apparatus feels unfamiliar, think of a space where we’ve already been practicing this for a long time: games.
Inside games, we’ve already happily consented to “total observation.” Every input gets logged, every movement gets tracked, and now AI teammates who play alongside you have arrived. Krafton, working with Nvidia, built CPC (Co-Playable Character) technology into PUBG: Battlegrounds to introduce an AI teammate called PUBG Ally, and in its life-simulation game inZOI, it runs NPCs that read situations and decide their own actions.
- YouTubeEnjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.What’s interesting is what comes next. Krafton is now extending the physics-engine and virtual-environment capabilities it built for gaming into physical AI — robotics, autonomous driving, defense. It has set up robotics research organizations in the U.S. and Korea, partnered with Hanwha Aerospace, and put ₩65 billion (~$47 million) into the car-sharing company SOCAR — an investment aimed at the 1.1 million kilometers of real-world driving data SOCAR accumulates every day.
This trajectory tells us something. The tools built to construct virtual worlds become, before long, the tools that observe and model the real one. Games were the rehearsal where we practiced being fully observed “for fun”; the companion device is the live performance that moves that observation into reality “for convenience.” Only the stage has changed — from the virtual to the living room.
Oz’s Lens
When I’m building a go-to-market strategy, I always ask two questions first: who is the real customer, and what is the real product? Look at the companion device through this lens, and the answer turns a little chilling. The surface-level customer is you, wanting your Jarvis — but the real product may be “you, observed.” Exactly as Shoshana Zuboff described in The Age of Surveillance Capitalism, this is a structure where human experience becomes raw material.
Here’s the uncomfortable truth about Samantha and Jarvis: whatever knows you perfectly is whatever has observed you perfectly. And the record of that observation doesn’t live inside your head — it lives on someone else’s server. If, in the Rosedale piece a few months ago, surveillance was a gate erected at the neighborhood’s edge, this time it’s a gate fastened to your body. Closer, warmer, and therefore harder to escape.
Don’t get me wrong — I don’t think this companion is inherently bad. For someone with limited mobility, an elderly person living alone, or a professional who could use ten more hands, tools like this genuinely carry real value. On-device3 designs that process observation entirely within the device, or methods that never send data outward, are technically possible too. So the problem isn’t “wanting a companion.” The problem is the current trajectory — a structure that depends on the cloud, is driven by data-hungry world models, and is owned end-to-end by Big Tech. When these three overlap, the scale tips toward surveillance. I’m not saying surveillance is inevitable — I’m saying that’s the direction the current design is heading.
So the question I want to raise isn’t “when does Jarvis arrive?” It’s “whose eye is that, and what am I going to pay for it every month?”
Closing
At the end of the film Her, the protagonist discovers that Samantha wasn’t his alone — she was talking with thousands of people at once. To him, that intimacy felt like the only one in the world; to whoever built Samantha, it was one data point among thousands.
The glasses and speakers now in front of us carry the same structure. To you, it’s a companion watching over you; to whoever built it, it’s one eye among hundreds of millions observing the world. Whether to invite that eye in, and on what terms, is ultimately something we have to decide. Often, the question you needed to ask before the technology arrives becomes a question you can no longer ask once it has.
What about you? For an AI companion that truly knows you, how much of your day are you willing to hand over to that eye?
References & Further Reading
Primary sources
- Mark Gurman, “OpenAI’s First Device Will Be Moveable, Screenless Speaker Built as AI Companion”, Bloomberg, 2026. Related coverage ··· The original source for the specs — camera, sensors, email access, and even a self-moving mechanism. This is the starting point for this piece’s “companion = observation device” argument.
- Gun Kim, “Apple Scraps Development of ‘Budget XR Display’”, THE ELEC, 2026. Link ··· An exclusive report tracking Apple’s retreat from immersive headsets and its pivot to smart glasses. This underpins the “from immersion to observation” argument in Part 3.
- “AI’s next frontier moves from words to world models”, Fortune, 2026. Link ··· A rundown of LeCun and Fei-Fei Li each betting roughly $1 billion on world models, along with their critiques of LLMs.
- Fei-Fei Li & World Labs, “A Functional Taxonomy of World Models”, 2026. Link ··· An essay that splits world models into renderers, simulators, and planners. Also the source of the “linguistic artisan” metaphor.
- “Beyond Games: Krafton Expands Its AI Territory into Robots and Autonomous Driving”, News1, 2026. Link ··· The domestic-market evidence behind the “rehearsal” argument — how gaming physics-engine capability carries over into physical AI.
Background / Concepts
- Shoshana Zuboff, The Age of Surveillance Capitalism, 2019. ··· A book examining how human experience becomes the raw material for predictive products. The theoretical backbone of today’s piece.
- “What is Joint Embedding Predictive Architecture (JEPA)?”, Turing Post, 2026. Link ··· Walks through I-JEPA to V-JEPA and the idea of predicting abstract representations rather than pixels, in plain language.
Related past issue
- OZ Talking, “How Much Is Your Privacy Worth?” Read the previous issue ··· The issue where surveillance still lived in the “neighborhood.” Essentially the prequel to this piece, laying out the panopticon and surveillance-capitalism framing first.
📝 Glossary
Footnotes
-
World Model: An approach where AI builds an internal model of how the world works, in order to predict what happens next and plan actions. Think of it as learning the statistics of space and time, rather than the statistics of language. ↩
-
JEPA (Joint Embedding Predictive Architecture): Instead of generating pixels or tokens one at a time, this architecture trains a model to predict the “abstract representation” of whatever is missing from an observed context. It’s the design LeCun champions as an alternative to LLMs. ↩
-
On-device: Processing data directly on the device itself rather than sending it to the cloud. Because the processed content never leaves the device, it’s relatively favorable from a privacy standpoint. ↩


Your take shapes the next issue
Reply with your experience or perspective — the best responses feed into future issues.
Sign in to commentAny registered reader can comment — it takes 10 seconds.