The general consensus emerging across the AI and industrial spheres is that the models themselves are no longer the bottleneck — they are more than good enough to power deep conversations, produce vast swathes of production-grade code, and resolve all manner of customer support requests.
The harder problem is everything wrapped around the models: the policies that decide what an agent can do on its own, the escalation path for when it can’t, and the record showing what happened when something goes wrong.
Ultimately, it boils down to trust
Ultimately, it boils down to trust: As a head of customer operations or a chief information officer, do you want a system handling live customer calls and taking actions on someone’s account without a clear boundary on what it’s allowed to handle unsupervised, and who’s accountable if it fails?
OpenAI thinks it has an answer in the form of Presence, a product unveiled on Wednesday that puts AI agents — the same ones OpenAI has been running on its own support line — on enterprise phone and chat channels.
In a blog post announcing the product, OpenAI notes that the real test for enterprise AI agents has moved beyond proving they can do the job — now it’s whether they can stay reliable as the products, policies, and people around them keep changing.
“The challenge for enterprises is no longer proving that AI agents can work; it’s making them reliable enough to do high-value work in production,” the company writes. “Agent behavior must also adapt as products, policies, and user behavior change.”
“The challenge for enterprises is no longer proving that AI agents can work, it’s making them reliable enough to do high-value work in production.”
Presence of mind
With Presence, a company picks a single, specific job for the agent to handle — an insurance claim, an IT request, a billing dispute — and the agent only gets the systems and information tied to that job, nothing more. It’s the company, not OpenAI, that sets the rules: what the agent can do without checking in, what needs a person’s sign-off first, and the point at which it stops and hands over to a human.
Getting to that point, however, isn’t a matter of flipping a switch on an API. OpenAI’s own engineers sit with each customer to figure out the job, wire up the systems, decide the permissions, run it through testing, and get it live, then hand ongoing support to outside integration partners as the deployment grows.
Today, Presence covers real-time voice and chat — customer support, outbound sales, higher-risk internal requests.
Notably, OpenAI says this has actually been part of its own internal support operation for some time, where it claims the agent now resolves 75% of inbound issues on its English-language phone line without a person stepping in. And that’s obviously a big part of its pitch: if it’s good enough for OpenAI, one of the most valuable private companies in the world, whose reputation depends entirely on people trusting what it builds, then it surely must be good enough for anyone else?
Whether it is, or isn’t, remains to be seen. But the company has a handful of early design partners on board, including BBVA, SoftBank, and IAG. And it’s clear this is still far from a full rollout — it’s not yet available as a self-serve product, which is one such signal. Access, for now, is limited to eligible enterprise customers through a restricted rollout, with deployments handled directly by OpenAI’s own “forward deployed engineers” and a small number of global systems integrators.
Presence isn’t OpenAI’s only move on this front. The launch lands a month after the ChatGPT-maker, alongside tech titans including Google and Microsoft, founded the Appia Foundation, a Linux Foundation effort to give companies a standardized way to demonstrate their AI systems meet safety and compliance obligations, as opposed to relying on self-declared claims.
Where Appia is building industry-wide paperwork for proving an AI system can be trusted, Presence is OpenAI trying to prove it one customer at a time.
The forward deployed engineer: AI’s trust layer
Zooming out, Presence’s dependence on OpenAI’s own engineers is itself part of a broader trend, with tech companies embedding technical staff directly with a customer to design, build, and support an AI system on-site rather than handing over an API and walking away.
As The New Stack has previously reported, the “forward deployed engineer” (FDE) exploded into one of the industry’s most sought-after jobs within the space of about ten days in May, with OpenAI launching its own $4 billion company built entirely around staffing enterprises with these engineers, and Google posting dozens of openings with salaries reaching into the mid-six-figures. AWS, meanwhile, announced in June that it was putting $1 billion behind a similar team of its own, embedding engineers directly with customers to help build and run AI systems using the customer’s own data and infrastructure.
“It’s easy to make an AI agent. The hard part is making an AI agent which can be trusted to speak directly with customers today and which will adapt as needs change.”
Colin Jarvis, OpenAI’s global head of forward deployed engineering, took to LinkedIn on Wednesday to explain that Presence grew out of OpenAI’s FDE team working directly with its own product team to solve a recurring problem in customer deployments.
“A lot of our most challenging customer work begins with external-facing use cases where agents need to be robust to changing customer behavior, business policy shifts and external compromise,” Jarvis writes.
What that points to is trust: a company handing an agent its phone lines wants to know a person who understands both the system and the business is standing behind it, not just a model running unsupervised.
Zach Parent, a forward deployed engineer at OpenAI, explains on LinkedIn that the difficulty has now shifted from building an agent at all, to building one a company will actually let near its customers as its needs keep changing.
“These days, it’s easy to make an AI agent,” Parent writes. “The hard part is making an AI agent which can be trusted to speak directly with customers today and which will adapt as needs change.”
Group Created with Sketch.
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.