Voice-First AI Support Was Designed for a Headset, Not a Hard Hat
This week OpenAI launched Presence, its first real enterprise product beyond the model layer: a platform for deploying AI agents that handle customer support and internal service requests over voice and chat, wired into company data, policies, and existing software. Two days later, TeamViewer and ServiceNow announced a multi-year partnership to fold TeamViewer’s digital employee experience and remote connectivity stack into the ServiceNow AI Platform, chasing “autonomous IT operations.”
Two of the biggest platform plays in enterprise support, in the same week. And both of them are built around the same customer: a person sitting at a desk, wearing a headset, looking at a screen.
Neither pitch has a visual layer anywhere in it. Voice agents. Chat agents. IT endpoints. Device telemetry. Not one word about the support platform actually seeing the thing that’s broken.
The desk worker is the easy case
Look at what Presence is designed to automate: billing issues, insurance claims, employee IT requests. Look at what TeamViewer + ServiceNow is designed to fix: laptops, software, digital workplaces. These are problems that live entirely inside computers. The AI can read the ticket, query the system, push the fix, and close the loop without anyone ever looking at anything physical.
That’s not a criticism. It’s a huge market and the automation genuinely works there. But it works because the problem is text-shaped — the entire state of the problem is already digital. The AI has perfect information because the information was never anywhere else.
Now take that same voice agent and put it in front of a field technician standing on a ladder next to a rooftop HVAC unit. Or a homeowner staring at a water heater that’s making a noise they can’t name. Or a factory operator looking at a fault code on a machine with forty nearly identical connectors behind an access panel.
What does “voice-first” get you there?
Describing a physical problem out loud is a terrible interface
Here’s the actual UX that voice-first AI offers the hard-hat worker: stand in front of the broken thing and try to compress a three-dimensional physical reality into spoken words, so a language model can decompress those words back into a guess about what you’re looking at.
“There’s a valve — no, more like a fitting — on the left side, kind of behind the copper pipe. It’s leaking from the top. I think. There’s corrosion, or maybe it’s just old flux.”
Every sentence is lossy. Every sentence is a chance for the customer to use the wrong word, and for the AI to confidently run with it. Language models are spectacular at processing descriptions. They are completely at the mercy of the description’s accuracy. A human support agent at least knows when a caller sounds unsure. An AI agent hears “the valve on the left” and treats it as ground truth.
The failure mode isn’t the AI misunderstanding. It’s the AI understanding perfectly — and being perfectly wrong, because the input was wrong. I wrote about this gap when the agentic support wave first hit: agentic support is here, but seeing the customer’s problem isn’t. Two major launches later, nothing has changed.
The hard hat is the stress test
The person in a hard hat is the stress test that voice-first support fails, for reasons that have nothing to do with model quality:
- Their hands are dirty or occupied. They can’t type into a chat window. Fine — voice handles that. But voice can’t handle the next part.
- Their problem is visual. Which wire. Which connector. Which error light. What color the corrosion is. Whether the seal is seated. None of this survives translation into speech.
- They don’t share vocabulary with the AI. The customer says “the metal box thing.” The knowledge base says “condensate pump assembly.” A camera resolves that mismatch in one second. A conversation never does.
- The environment is loud. Machine rooms, job sites, roadside. Voice recognition degrades exactly where physical problems concentrate.
Meanwhile, every one of these people is holding a device with a better sensor than their vocabulary: a smartphone camera. The information the AI desperately needs is one camera activation away, and the entire industry keeps building elaborate systems to have people describe it instead.
Vision is the missing primitive in the platform race
Here’s what makes this strange: the models can already see. Multimodal AI is arguably the most impressive capability shipped in the last two years. GPT-class models can identify a part from a photo, read a fault code off a blurry nameplate, and spot an installation error a first-year tech would miss.
The capability exists. What’s missing is the plumbing — the visual layer in the support platform itself. The channel that gets live video from a customer’s or technician’s phone into the support workflow with zero friction, so a human agent or an AI agent can look at reality instead of a transcription of it.
OpenAI built agent-to-data plumbing. TeamViewer and ServiceNow built agent-to-endpoint plumbing. Nobody in either announcement built agent-to-eyeball plumbing. The platform race is being run entirely on problems that are already digital, while the expensive problems — truck rolls, repeat dispatches, returned parts with no fault found, escalations that ping-pong for days — are physical.
Every support economics conversation eventually lands on the same numbers: a truck roll costs hundreds of dollars, a misdiagnosed dispatch costs two truck rolls, and the difference between them is usually thirty seconds of accurate visual information at first contact.
Point the camera at it
The fix isn’t exotic. When the problem is physical, the first move should be to see it — not to interrogate the customer about it. That’s the entire premise behind Viewabo: the agent sends a link, the customer taps it, and their smartphone camera becomes the agent’s eyes. No app download, no account, no “can you describe where the leak is coming from.” The support session starts from shared visual ground truth instead of a game of twenty questions.
And once the visual channel exists, it compounds with everything the platform vendors are building. An AI agent with access to live video is dramatically more useful than one parsing verbal descriptions. Visual context turns the voice agent from a guesser into a diagnostician. The visual layer doesn’t compete with the Presence-style agent platforms — it’s the input those platforms are missing for the entire physical-world half of support.
The headset worker is well served. The platforms launched this week will serve them even better. But support doesn’t stop at the edge of the desk, and the next platform battle will be won by whoever remembers that the most valuable sensor in customer support isn’t a microphone.
It’s the camera the customer is already holding.
