Agentic Support Is Here. Seeing the Customer’s Problem Isn’t.
July 22 was the day enterprise agentic support stopped being a roadmap slide.
In one news cycle, OpenAI launched Presence: a fully managed platform that deploys voice and chat AI agents into enterprise workflows. On the same day, ServiceNow announced its AI business crossed $1 billion in annual contract value. Customers running agentic AI in production grew 9x in nine months.
Read those two announcements together and the message is unambiguous: AI agents answering customers are no longer an experiment. They’re a billion-dollar production category, and the biggest players in software are now selling the deployment muscle to make them stick.
And yet both announcements share the same blind spot. Literally.
What Actually Launched
Start with Presence. This isn’t another API. OpenAI is explicitly not handing you a toolkit and wishing you luck. Presence is a managed platform: you work with OpenAI’s forward-deployed engineers and systems integrators to put production-ready voice and chat agents into your existing workflows, complete with guardrails, policy controls, and a continuous improvement loop. There’s no self-serve signup. It’s enterprise deployment as a service.
That design choice tells you something important. OpenAI looked at the wreckage of DIY agent deployments — the 70% production failure rates, the Gartner cancellation projections — and concluded that the missing ingredient was hands-on deployment expertise. They’re probably right, as far as it goes.
ServiceNow’s numbers tell the demand side of the story. AI ACV past $1 billion. AI net-new ACV up more than 40% sequentially. Customers running agentic AI in production up ninefold in nine months. That’s not pilot budget. That’s committed enterprise spend on agents doing real work, right now.
So the “is agentic support real?” debate is over. It’s real, it’s funded, and it’s scaling faster than almost any enterprise software category in memory.
Here’s what nobody launched on July 22: an answer for what happens when the customer’s problem can’t be typed or spoken.
Text and Voice Are Two-Thirds of a Support Stack
Look at what Presence actually handles: real-time voice and chat. Look at what ServiceNow’s agents do: triage tickets, resolve IT requests, automate workflows — all mediated through text. The entire enterprise agentic stack, from the frontier lab to the workflow giant, is built on two modalities: what customers say and what customers write.
That’s fine when the problem lives in a database. Reset a password, check an order status, process a return, update a record — agents will eat that work, and they should.
But an enormous slice of support doesn’t live in a database. It lives on a kitchen counter, in a server closet, on a factory floor, behind a router with a blinking light the customer can’t name. Physical products. Hardware setups. Error states on device screens. Installations that are “done” according to the customer and visibly wrong according to reality.
Where Agentic Support Goes Blind
For that slice, the world’s most sophisticated voice agent is in exactly the same position as a 2005 call center rep: taking the customer’s word for it.
- “The cable is plugged in.” Is it? Into the right port?
- “I already restarted it.” Did the restart complete, or did they just toggle the screen?
- “The light is red-ish, sort of orange.” Which of the four lights?
An agent can’t verify any of this. It can only infer from descriptions provided by the least qualified describer in the interaction — the customer, who by definition doesn’t understand the product well enough to fix it themselves. Garbage description in, confident wrong answer out. And agents are trained to sound confident.
This is the pattern I keep coming back to: agents fail at the edges of their information environment. I wrote about it last week — what 830 IT leaders are getting wrong about agentic AI in support — and the July 22 news makes it more urgent, not less. Enterprises are scaling text-and-voice agents 9x while the perception gap that breaks those agents remains completely unaddressed.
Why the Big Platforms Won’t Fix This First
You might assume OpenAI or ServiceNow will just bolt on vision. The models can already interpret images. What’s the holdup?
The holdup is that model capability isn’t the bottleneck — the interaction is. Getting live visual context from a customer isn’t a model problem. It’s a workflow problem:
- You need a frictionless way to get a camera stream from someone’s phone into the support session. No app downloads, no account creation, no IT approval on the customer’s side.
- It has to work at the exact moment of frustration, in one tap, or the customer bails.
- And the human agent (or the AI, when it’s ready) has to see, annotate, capture, and record. That visual record must attach to the ticket so escalation doesn’t restart from zero.
That’s not a feature you ship in a quarter by fine-tuning a model. It’s product infrastructure that has to sit in the messy seam between the customer’s pocket and the enterprise’s workflow. The big platforms are busy winning the text-and-voice land grab because that’s where the billion-dollar ACV is today. Visual context is somebody else’s problem.
It’s ours, actually. Making the customer’s camera a first-class support channel — no app, one link, live video into the agent’s workflow — is exactly what Viewabo does. Not as a bolt-on, but as the missing modality.
The Strategic Read for Support Leaders
If you run support for anything with a physical footprint, here’s how to interpret July 22:
Deploy agents aggressively for the text-shaped work. The economics are real, the platforms are maturing, and the deployment help now exists. Holding out is a losing position.
But audit your ticket mix honestly. What percentage of your issues require verifying real-world state — a device, a connection, an installation, an error screen? For hardware, IoT, appliances, networking, and field service, that number is usually far higher than leadership assumes. Why? Those tickets hide inside categories labeled “user error” and “no fault found.”
Then design the visual path before the agent hits the wall. When an agent reaches its perception boundary, the escalation shouldn’t be a phone call where a human starts the same blind guessing game. It should be a one-tap jump to live video. The customer shows the problem, and the human (with the AI transcribing alongside) sees it. The ticket resolves in minutes instead of days — or instead of a truck roll.
Agentic support is going to sort organizations into two groups: those whose AI can only process what customers claim, and those whose support stack can verify what’s actually true. Text and voice got their billion-dollar coming-out party on July 22.
Sight is still up for grabs. The companies that add it first will hold the advantage every blindfolded agent in the industry can’t touch: they’ll be the only ones who can actually see the problem.
