OpenAI Presence Explained [How Enterprise Agents Work]
OpenAI Presence is a deployment product for enterprise agents, not a new model, API, or do-it-yourself agent builder. It combines a model with a narrowly defined job, the minimum systems and knowledge required for that job, approved actions, policy constraints, human escalation, simulations, graders, and a Codex-assisted improvement loop. OpenAI and the customer implement each deployment together, with field-deployed engineers or select integrators, rather than handing an administrator a self-serve console.
That distinction matters. A model generates or reasons. An API exposes model capabilities to software. A builder helps a team assemble its own agent. OpenAI Presence packages a governed operating deployment for one enterprise job. OpenAI Frontier remains the broader enterprise platform and program around agents; Presence is the deployed-agent product announced inside that strategy.
This guide reflects published information as of July 22, 2026. Presence entered limited general availability that day, so the launch evidence is specific but early. It does not establish independent cross-industry performance, pricing, compliance eligibility, or staffing impact.
TL;DR: Presence packages a single enterprise job, bounded permissions, tested escalation, and reviewed improvement inside a managed deployment—not a general license for an agent to act anywhere.
Key Takeaways
Presence deploys one scoped job at a time with the minimum knowledge and system access needed to perform it.
Policies define permitted behavior; approved actions constrain what the agent can do; human escalation handles exceptions and higher-risk decisions.
Simulations and graders test the workflow, while Codex can propose controlled improvements that still require review and rollout discipline.
OpenAI reports results from its own English-language support line, not an independent benchmark for every company or industry.
Limited general availability is for eligible enterprises working with OpenAI field-deployed engineers or select integrators; it is not a self-serve product.
The practical buying question is whether a measurable, bounded workflow can justify the integration and governance work—not whether a general chatbot sounds capable.
OpenAI Presence in One Sentence
OpenAI Presence is a jointly deployed enterprise-agent product that gives a model one bounded job, limited context and tools, explicit policies and approved actions, tested escalation paths, and a reviewed Codex improvement loop; unlike a model, API, builder, or the broader Frontier program, it arrives as an operated workflow.
What OpenAI Actually Announced
OpenAI described a deployment pattern rather than a new foundation model. Each agent receives one job, access to only the knowledge and systems needed for that job, rules governing what it may say or do, a catalog of approved actions, and criteria for when a human takes over. Voice and chat are the initial customer-facing channels, with customer support, outbound sales, and high-risk internal operations named as use cases.
According to OpenAI and corroborating launch coverage from Reworked, its own English-language phone-support deployment resolved 75% of conversations without a human and reduced handoffs by 15 percentage points in 10 days. Those are vendor-reported results from one OpenAI workflow, not an independent Presence benchmark.
OpenAI reports a 75% resolution rate on its own English support line. The OpenAI announcement and fetchable Reworked report tie that figure to one disclosed workflow; it should not be transplanted into another organization’s forecast.
Handoffs fell 15 percentage points over 10 days in OpenAI’s deployment. Both OpenAI and Reworked preserve that scope; the short window makes denominator, repeat-contact, and resolution definitions important.
OpenAI also names enterprises at different exploratory stages. BBVA is exploring the product, SoftBank is testing it, and IAG is exploring a severe-weather use case. None of those descriptions is evidence of a production result. That language should remain intact when evaluating customer fit.
Presence Versus Adjacent OpenAI Products
| Offering | Primary unit | Who assembles the workflow | Governance included | Deployment status |
|---|---|---|---|---|
| Foundation model | Model response or reasoning step | Customer or software vendor | Model-level safeguards | Available by model terms |
| API | Programmatic model call | Customer engineering team | API controls, not a finished workflow | Self-serve or contracted |
| Agent builder | Agent configuration | Customer team | Builder-dependent controls | Customer-operated |
| OpenAI Frontier | Enterprise agent platform and program | Enterprise with OpenAI ecosystem | Broader identity, data, and agent management | Enterprise program |
| OpenAI Presence | One deployed, governed job | OpenAI plus customer or select integrator | Policies, approved actions, escalation, simulations, graders | Limited GA |
Sources: OpenAI’s Presence announcement and Reworked’s launch analysis.
This is the core differentiation: Presence is closer to a production operating model than a toolkit. It still depends on models, APIs, enterprise systems, and human owners, but the product boundary includes implementation and improvement rather than stopping at access to technology.
A Machine-Readable Capability Map
| Capability | Required input | Produced output | Primary control | Human boundary |
|---|---|---|---|---|
| Understand the request | Voice or chat turn | Intent and context | Allowed job scope | Ambiguous identity or intent |
| Retrieve enterprise knowledge | Approved sources | Grounded answer context | Minimum data access | Missing or conflicting source |
| Take an action | Authorized system state | Approved transaction | Action allowlist and policy | Irreversible or high-risk action |
| Escalate | Trigger or low confidence | Routed case with context | Escalation rule | Human accepts ownership |
| Test behavior | Simulated conversations | Grader results | Pass/fail thresholds | Reviewer approves release |
| Improve deployment | Production evidence | Proposed workflow change | Codex-assisted change review | Owner approves controlled rollout |
The table is deliberately operational. “Agent” can otherwise collapse several separate decisions into one label. A reliable deployment defines the input, output, control, and human boundary for every capability. An agent that can retrieve a policy but cannot prove identity should not be allowed to alter an account. An agent that can draft a response may still need a person to approve the send.
Availability: What Limited GA Means
| Availability fact | Published value | Date or count | What it does not mean |
|---|---|---|---|
| Launch date | Limited GA | 07/22/2026 | Broad self-serve access |
| Eligible route | Enterprise engagement | 1 scoped deployment path | Instant workspace activation |
| Implementation parties | OpenAI FDE or select integrator | 2 partner routes | Any vendor is authorized |
| Named channels | Voice and chat | 2 channels | Every channel is supported |
| Named enterprise examples | BBVA, SoftBank, IAG | 3 organizations | 3 production deployments |
| OpenAI support evidence | English phone line | 1 disclosed workflow | Cross-industry validation |
Sources: OpenAI and the fetchable Reworked report.
| Launch evidence | Published figure | Scope figure |
|---|---|---|
| Automated resolution | 75% | 1 disclosed workflow |
| Handoff movement | -15 pp | 10 days |
| Initial channels | 2 | 2 channel types |
| Named enterprise examples | 3 | 3 organizations |
| Availability route | 1 limited-GA product | 2 implementation paths |
According to Reworked, the July 22, 2026 release is limited to eligible enterprises and is not self-serve; the article also distinguishes 3 named enterprise examples from proven production results. Procurement teams should therefore validate access, integrator responsibility, and support boundaries before designing a business case around the product.
The Resolution Metric Box
Resolution rate = conversations fully resolved by the agent ÷ conversations handled by the agent × 100. Define abandonment, repeat contact, reopened cases, and transferred conversations before comparing two systems. Track customer outcome and policy adherence beside containment so a lower handoff rate cannot hide a bad resolution.
According to Zendesk, a simple example with 150 fully resolved conversations out of 200 handled produces a 75% automated-resolution rate. Zendesk also warns that teams need explicit treatment of abandoned, unresolved, contained, and repeat contacts. That denominator discipline is more portable than any vendor’s headline number.
The operational scorecard should therefore hold at least five fields: eligible conversations, agent-handled conversations, resolved conversations, repeat contacts inside a defined window, and human handoffs. Add policy violations and unauthorized-action attempts even when their desired value is zero. A team can then separate “the agent stayed in the conversation” from “the customer’s job was actually completed.”
What Separate Production Deployment Evidence Adds
Presence itself has no independent cross-industry benchmark at launch. A separate production study is still useful for understanding how much context changes results—as long as it is not presented as Presence performance.
According to Nubank’s KDD production study, the company has more than 100 million customers, and its researchers analyzed 5 production deployments. One card-delivery deployment reported a 37-percentage-point increase in transactional net promoter score and a 29-percentage-point increase in self-service, but those figures describe Nubank’s architecture and use cases, not OpenAI Presence.
Nubank reports 100M+ customers company-wide and analyzes 5 production deployments. The company-authored research paper makes its different architecture and use cases explicit; its value here is methodological.
| Evidence item | Sample or period | Reported figure | Transferability to Presence |
|---|---|---|---|
| OpenAI English support | 1 workflow | 75% resolution | Direct product evidence, narrow context |
| OpenAI handoff change | 10 days | -15 percentage points | Direct product evidence, short window |
| Nubank company and study scale | 5 deployments | 100M+ company customers | Company-authored agent evidence, different stack |
| Nubank card delivery | 1 use case | +37 pp tNPS | Use-case evidence, not Presence |
| Nubank self-service | 1 use case | +29 pp | Use-case evidence, not Presence |
Sources: OpenAI Presence, Reworked, and the Nubank production study.
According to PwC, it announced an OpenAI-based agentic contact-and-service solution and a dedicated center of excellence on July 15, 2026. That provides ecosystem context, not proof that the PwC offer is Presence or that the two products share identical controls.
A Seven-Step Presence Deployment Playbook
Name one completed customer job. Use an outcome such as “confirm claim status after identity verification,” not a department-wide ambition such as “automate service.”
Set the eligible population. Define channels, languages, customer types, hours, and exclusions. Freeze this denominator before measuring resolution.
Map minimum access. List each knowledge source, read permission, write permission, and retained data field. Remove access that the job does not require.
Create the action allowlist. Separate read, draft, reversible write, irreversible write, and prohibited actions. Require approval at the appropriate boundary.
Write escalation contracts. Specify the trigger, destination, context bundle, response expectation, and ownership transfer. Test unavailable-human conditions too.
Run simulations and graders. Include ordinary requests, identity failures, conflicting policies, missing records, adversarial phrasing, and system downtime. Treat each change as a regression candidate.
Release, observe, and improve. Start with a bounded cohort, review outcome and control metrics, and let Codex propose changes only through a versioned review and controlled rollout.
At US Tech Automations, this becomes a concrete workflow artifact: the team maps each request trigger, retrieves only approved records, routes policy exceptions, logs every action, and hands a complete context packet to a human owner. The useful deliverable is not a generic “AI strategy”; it is a testable job definition with named systems, permissions, events, and failure handling.
Where Presence Fits by Industry
The same deployment pattern changes shape when the risk boundary changes. An insurance agent can collect first-notice-of-loss details and route an adjuster without deciding coverage. A healthcare agent can handle scheduling and billing status without interpreting symptoms. A property-management agent can open a work order and page an emergency contact without interpreting a lease or deciding an accommodation.
For the workflow-specific versions, see:
The control pattern also connects to earlier Frontier analysis. Verified Intelligence focuses on verification, simulation, and auditability, while Microsoft Service Agent shows a different product boundary for service workflows. Comparing boundaries is more useful than comparing brand labels.
Signal vs Speculation
Sourced signal: Presence launched in limited general availability; it is jointly deployed for one scoped job; it uses minimum access, approved actions, policies, escalation, simulations, graders, and a Codex-assisted improvement loop. OpenAI disclosed one internal support result and three named enterprise relationships at exploration or testing stages.
Speculation to test, not treat as fact: Presence may become a template for outcome-priced or deeply operated enterprise agents, and its improvement loop may compress the time between production evidence and a reviewed change. It may also create a clearer services market for integrators that can document permissions and escalation. None of those outcomes, along with pricing, staffing savings, compliance status, or broad production performance, was established in the launch material.
Frequently Asked Questions
Is OpenAI Presence a new AI model?
No. Presence uses models inside a deployed enterprise workflow. The product boundary includes a scoped job, enterprise context, authorized actions, escalation, testing, and improvement. A model is one component of that system.
How is Presence different from an API or agent builder?
An API provides programmable model access, and a builder helps a customer assemble an agent. Presence is implemented with OpenAI and the customer, or a select integrator, around one operated job. It is not a self-serve configuration surface at launch.
Who can buy OpenAI Presence now?
OpenAI says eligible enterprise customers can engage through its field-deployed engineering organization or select integrators during limited general availability. The announcement does not publish a universal eligibility checklist or price card.
Does the 75% resolution figure predict another company’s result?
No. It describes OpenAI’s own English-language phone-support deployment. Another organization needs its own eligibility rules, resolution definition, baseline, repeat-contact window, and controlled pilot.
Which decisions should always escalate to a person?
The answer depends on the workflow, but identity ambiguity, conflicting source records, policy exceptions, irreversible actions, regulated judgment, safety issues, and customer disputes are common escalation boundaries. Each trigger needs a named destination and context packet.
Can Codex change a live Presence agent automatically?
The announcement describes Codex proposing improvements inside a controlled loop. A sound operating model versions the proposal, runs simulations and graders, requires owner review, and releases through a bounded rollout rather than treating generated code as automatic authorization.
The Practical Decision
Presence is most relevant when a company can name one high-volume job, measure completion cleanly, restrict the data and actions involved, and staff the human exception path. It is less relevant when the desired outcome is vague, the source systems are inconsistent, or the organization has not decided who owns policy and production changes.
US Tech Automations can turn that decision into an implementation map: inventory the trigger and systems, define read/write permissions, build approval and escalation routes, instrument resolution and repeat contact, and stage regression-tested changes. If the job is bounded enough to deploy, explore the agentic workflow architecture that connects those controls without assuming a vendor benchmark will transfer.
About the Author

Helping businesses leverage automation for operational efficiency.
Related Articles
See how AI agents fit your team
US Tech Automations builds and runs the AI agents that handle this work end to end, so your team doesn't have to.
View pricing & plans