OpenAI announced Presence on July 22: an enterprise platform for deploying and managing AI agents across customer-facing and internal workflows. Early customers include BBVA, running voice support for routine banking in Mexico; SoftBank, handling natural Japanese-language customer conversations; and Australian insurer IAG, adding support capacity during high-demand periods. The pitch is not a smarter model. It is that deployed agents should be run like software: governed by policies, tested in simulation, scored by graders, and monitored in production. For operators, that framing matters more than the launch itself.
Key takeaways
- Agents as managed software. Presence packages policy and guardrail configuration, simulation testing against edge cases, grader-based evaluation, and production monitoring into one platform.
- High-volume support comes first. BBVA (voice banking in Mexico), SoftBank (Japanese-language conversations), and IAG (surge support) are the early deployments.
- Vendor numbers, vendor channel. OpenAI says 75% of inbound issues on its own English phone-support line resolve without human assistance. Treat that as a claim, not a benchmark.
- Not self-service. Presence ships through hands-on engagements led by OpenAI's engineers and partner integrators, with pricing undisclosed. Plan for services scope from day one.
- Do the commercial diligence early. Services cost and lock-in questions belong in the evaluation, before the pilot starts.
What OpenAI Presence does
Presence deploys agents that answer questions, use company systems, take approved actions, and escalate to humans under company-defined policies. Around that core sits the operational scaffolding: guardrail configuration, simulation testing against edge cases, grader-based evaluation of changes, and monitoring once agents are live.
There is also a continuous-improvement loop built on Codex, with controlled rollouts. OpenAI says that loop cut human handoffs on its own support channel by 15 percentage points over 10 days. The number is the vendor's, measured on the vendor's own workload, but the mechanism is the interesting part: changes ship through graders and staged rollouts, not through someone editing a prompt in production.
The industry is converging on acceptance criteria
Strip away the branding and Presence is an argument about process. Agents do not ship on demo quality. They ship against defined policies, get tested in simulation before release, get scored on every change, and get watched in production. That is acceptance-criteria discipline, sold as a product.
We have made this argument for a while: the agent pilot that succeeds is the one where pass/fail was written down before the build started. When the model vendor itself ships graders and simulation suites as the core of its enterprise offering, that stops being a consultant's preference and becomes the industry default. Whatever platform you evaluate, this toolchain is now the baseline to demand.
The evaluation harness just became the product. If you have not written your acceptance criteria, someone else's platform will write them for you.
A services engagement, not a product trial
Presence is a limited general-availability program, not self-service. Deployments are led by OpenAI Forward Deployed Engineers and select global systems integrators, with Bain & Company named as a partner. Pricing was not disclosed.
That shape changes what an evaluation looks like. This is not a SaaS trial you can spin up and quietly abandon. It is a services engagement with real integration work, and it puts OpenAI head-to-head with the contact-center platforms and BPO vendors operators already pay. Three questions follow directly:
- What does it cost, in full? Model usage plus engineer and integrator time. Get the services scope and pricing in writing before committing pilot resources.
- Who owns the artifacts? Policies, graders, simulation suites, and integration code are the real deliverables. If they only run inside Presence, switching costs compound with every improvement cycle.
- What happens at renewal? An agent handling a large share of inbound volume is operational infrastructure. Negotiate exit and portability terms while you still have leverage.
What to do before you pilot
If Presence, or any managed agent platform, is on your shortlist, the work starts before the vendor call.
- Write acceptance criteria first. Define resolution rate, escalation accuracy, and policy-violation thresholds on your own terms. Then run the vendor's graders against your criteria, not theirs.
- Price the whole engagement. Integration scope, engineer involvement, and ongoing services cost, in writing. Compare against your current cost per contact, not against zero.
- Benchmark against incumbents. Your contact-center and BPO vendors are shipping agent features too. Run the comparison on your own numbers.
- Start narrow. One workflow, clear escalation paths, measured against the criteria you wrote. Expand on evidence, not on momentum.
We have argued that agent programs stand or fall on acceptance criteria, not demos, and that the build-versus-buy decision deserves the same rigor before any platform commitment. If you are weighing a managed offering like Presence against building on your own stack, we can help you scope the comparison — book a consult.