What Autonomous Resolution Looks Like When the Problem Has a Physical Form

“Autonomous resolution” is the phrase of the year in customer support. Every vendor deck has it. Encore AI just raised $30 million from Team8 to train voice agents on your company’s call history. The goal: agents that handle support and sales interactions on their own. The pitch is seductive: the agent learned from your best reps, it runs their playbooks, it even tells their jokes. Interaction goes in, resolution comes out, no human required.

I want to take the phrase seriously — more seriously than the people selling it do. Because “autonomous resolution” has a precise meaning. And almost nobody using it has checked whether their product delivers when the problem isn’t made of data.

Resolution means the world changed

A resolution is not a completed conversation. It’s a changed state. The refund posted. A password works again. Your account unlocks. Notice what those examples have in common: the problem and its fix both live inside a database. The AI agent has API access to the system where the problem exists. So “autonomous” is literal — the agent can reach the broken state and change it.

Now put a furnace in the sentence. A router. A CNC machine flashing a fault code on a factory floor. The problem’s state doesn’t live in any system your agent can query. Instead, it lives in physical reality. A loose wire, a clogged filter, a valve turned the wrong way, an error screen the customer is squinting at. No API call changes that state. The industry’s definition of autonomous resolution quietly assumes the problem is text-shaped. I’ve written before about how the pay-per-resolution pricing model makes the same assumption. The definition isn’t wrong. Someone just scoped it, silently, to the easy half of support.

What resolution requires when the problem is physical

Strip it down and resolving a physical-form problem requires three capabilities, in order:

See the state. Before anything else, something has to observe the actual condition of the actual object. Not the customer’s paraphrase of it — the thing itself. Model numbers, indicator lights, cable routing, corrosion, the error code as displayed rather than as remembered. Every diagnosis downstream is only as good as this observation.

Diagnose from visual evidence. Reasoning over what’s seen: this light pattern plus that sound means the igniter, not the gas valve. This is where AI is genuinely getting good, fast. Vision models can now read a nameplate and spot a miswired terminal. They can match a fault state against a corpus of resolved cases better than a first-year tech.

Change physical reality. Someone or something has to flip the breaker, reseat the cable, replace the filter. Unless you’re shipping a robot to every customer’s basement — you’re not — this step runs through human hands. The customer’s hands, or a technician’s.

That third step is the wall. An agent that can’t reach it hasn’t resolved anything; it has produced a well-written suggestion. And a suggestion delivered blind, without step one, is worse than nothing. It sends people to reset routers that were never the problem.

The realistic autonomy stack

So what can vendors honestly automate for physical problems? More than skeptics think, less than the decks claim. Here’s the stack as I’d actually build it:

Intake and triage: fully autonomous. AI should own the front door completely. Classify the issue, pull the account, check entitlements, decide whether this is text-shaped or physical-shaped. No human needed.

Observation: machine sight through the customer’s camera. The moment the problem is physical, the pipeline needs eyes. The cheapest, fastest sensor deployment in history is the smartphone already in your customer’s pocket. A live video session turns “it’s making a weird noise” into an observable, analyzable state. No app install, no truck.

Diagnosis: AI-led, human-checked at the margins. Vision plus reasoning over live video will handle a growing share of diagnoses outright. The escalation path to a human expert stays, but it gets thinner every quarter.

Actuation: guided human hands. This is the honest definition of “autonomous” in the physical world. Not the absence of humans, but the absence of dispatched humans. The system sees, diagnoses, and then guides the hands already on-site through the fix. It corrects in real time when the wrong screw comes out. The customer, in other words, becomes the actuator. And when the fix genuinely needs a licensed tech? The same visual record means the truck rolls once, with the right part.

That’s a credible autonomous-resolution pipeline for physical problems. Anything sold as autonomous without a visual channel has simply defined away the hard half of the job. It resolves the tickets that were always going to be cheap and punts the rest to a dispatch queue. I made this argument when agentic support platforms first started claiming end-to-end resolution.

Ask vendors one question

The timing of this definitional sloppiness is bad, because enterprises are deploying these agents faster than they can evaluate them. SAP just flagged “agent sprawl” as a board-level governance issue. Their LeanIX survey found 98% of companies have deployed AI agents or plan to. When adoption outruns scrutiny, marketing definitions become de facto architecture. Buy an “autonomous resolution” platform that can’t see, and you’ve architected a system that autonomously resolves password resets. Meanwhile, your furnace calls, router calls, and machine-down calls pile up in the same queue as before. Now there’s an extra bot in front of them.

So ask the one question that sorts the category instantly. When the problem has a physical form, how does your agent observe its state? If the answer is “the customer describes it,” you’re buying a chatbot with a thesaurus. If the answer involves a camera, you’re looking at a real pipeline.

This is where remote visual support slots in. It’s not a feature bolted onto a ticketing tool, but the enabling layer between diagnosis and actuation. It’s the piece that makes “autonomous” honest for the half of support that lives in the physical world. At Viewabo we’ve built our whole product on that layer. I’ve made the case that it should be the first architectural decision in a support stack, not the last. The vendors raising rounds on autonomous resolution will get there eventually. The definition will force them to. Physical reality doesn’t negotiate with a language model.