SAN FRANCISCO, 23 AUG 2026 — OpenAI has released Presence, a platform for enterprises to deploy AI agents that answer questions, use company systems, take approved actions and escalate to a person when required. It runs real-time voice and chat.

The product is not generally available. Deployments are limited to eligible customers and handled by OpenAI's own engineers alongside selected systems integrators — a delivery model that describes the state of agentic software more accurately than any capability list.

What Presence packages

Voice and chatReal-time, both channels
Gated rolloutEligible customers, OpenAI engineers and chosen integrators
BBVA, SoftBank, IAGAmong those exploring or testing it
Knowledge, procedures, actions, simulations, guardrails, escalationWhat the platform bundles

Named tasks include resolving billing disputes, supporting insurance claims, handling employee IT requests and conducting outbound sales. Customers determine which systems and information an agent may reach, what actions it may complete, and when human approval is required.

Spain's BBVA is evaluating it for everyday banking support in Mexico, SoftBank is running trials on Japanese-language customer interactions, and Australian insurer IAG is assessing whether it can absorb demand surges during severe weather.

The gated rollout is the most informative part of the announcement

OpenAI sells self-serve products. Anyone with a card can use the API. Presence is deliberately not that, and the reason matters more than the restriction.

An agent that resolves a billing dispute has to reach the billing system, understand which adjustments are permissible, apply the organisation's own rules, and know when to stop. None of that ships in a model. It is integration work against systems that differ at every customer, encoded in procedures that are frequently undocumented because the humans doing the job absorbed them by experience.

Requiring your own engineers on every deployment is what a vendor does when the product cannot yet be handed over. It is an honest position, but it is also a constraint. A business gated on the availability of vendor staff scales like a consultancy, not like software.

The optimistic reading is that this is temporary, and that each deployment produces reusable patterns until self-serve becomes possible. The sceptical reading is that enterprise process is irreducibly bespoke and the services requirement never goes away. Both are consistent with what has been announced.

The named tasks share one property

Billing disputes, insurance claims, IT requests and outbound sales look like a list of use cases. They are better understood as a single category.

All these tasks are high-volume, procedurally bounded, and currently performed by people following rules they did not write. They have defined, checkable outcomes. And they are already measured to death, giving the organisation a clear baseline for handling time and resolution rate to judge an agent against.

That measurability is why these tasks come first. They are not the most valuable work in an enterprise, but they are the work where a vendor can prove the product functions. Tasks with contested outcomes or no baseline will come much later, if at all.

Escalation is where these systems will actually be judged

The bundle includes escalation rules, and that component deserves more scrutiny than the conversational quality everyone will focus on.

An agent that handles the routine eighty per cent and escalates the rest changes the composition of the human queue rather than shrinking the work uniformly. What reaches a person is what the agent could not resolve, which is by construction the harder, angrier and more unusual cases.

Organisations consistently fail to plan for two consequences. First, handling time per human-touched case rises because the easy calls are gone; any productivity target built on the old average will be missed. Second, the humans remaining need to be more skilled, not less — the opposite of the usual staffing assumption for these deployments.

The other risk is escalation that fails to trigger. An agent that does not recognise it is out of its depth produces a confidently wrong resolution to a claim, and in insurance or banking that is a regulated outcome rather than a customer service problem.

Why the regional test cases are the interesting ones

SoftBank's trial of Japanese-language interactions is a more demanding test than any English deployment and, for that reason, worth watching.

Japanese customer service operates with register conventions, honorific levels and indirectness norms that are not stylistic decoration — using the wrong level is a substantive error that customers notice immediately. Model quality in a language is not the same as competence in that language's service conventions, and the second is harder to evaluate and harder to fix.

For contact centres across ASEAN the relevant question is the same one in a more complicated form. Operations here routinely handle several languages, frequent code-switching within a single conversation, and regional variants that training data underrepresents. An agent that performs well in English and poorly in Bahasa Indonesia does not reduce headcount; it segments the queue by language and leaves the harder segment fully staffed.

The region also hosts a substantial outsourced contact-centre industry, particularly in the Philippines, which makes the employment question here concrete rather than abstract.

The jobs question is being asked badly

Coverage of Presence has reached the employment question quickly, and mostly through the wrong frame.

Whether an agent replaces a person is the wrong unit of analysis, because contact centre work has never been staffed to headcount so much as to peak demand. IAG's stated interest is instructive: it wants to absorb surges during severe weather, which is capacity that currently either does not exist or is bought at short notice at high cost.

An organisation that uses agents to handle peaks without hiring temporary staff has changed its cost structure without making anyone redundant. That is a different outcome from replacing a standing team. It is also the harder sell internally, because the savings are invisible in headcount and appear only in the volume that would otherwise have been abandoned.

Where displacement does occur it will show up first in outsourced contract renewals rather than in direct redundancies, because that is where the flexible capacity sits. For the Philippines in particular, the contract cycle is the place to watch, not employer announcements.

What remains unconfirmed

Pricing is not described in the material reviewed, nor is what qualifies a customer as eligible. No deployment scale, resolution rate or accuracy figure has been published, and the named companies are described as exploring or testing rather than as having deployed in production.

Key details are missing. The announcement does not state which languages are supported or to what quality. It leaves undescribed how guardrails and simulations work, what happens when an agent takes an incorrect approved action, where liability sits, and what data leaves the customer environment. There is no mention of whether self-serve availability is planned.

What to watch for

The first signal is a customer moving from trial to production and saying so with numbers. Enterprise pilots of conversational systems have a long history of not converting, and the trial-to-production ratio is the only real evidence of whether this works.

The second is whether the services requirement relaxes. Self-serve Presence would mean the integration problem has been genuinely solved rather than absorbed by staffing.

The third is any published escalation rate. A vendor willing to say what proportion of conversations reach a human, and how that changed handling time for those cases, is describing the actual economics. Everyone else is describing a demo.