1. ‘Available today’ does not mean click-to-buy

OpenAI announced Presence on July 22, 2026 for voice and chat agents. Its Help Center narrows that availability: limited general availability for eligible enterprises, delivered by Forward Deployed Engineers and selected integrators, with no self-serve product.

There is no public price card or fixed integration catalogue. Workflow fit, implementation readiness and delivery capacity are established during scoping. This is the launch of a managed buying and deployment path, not mass-market SaaS.

2. What is actually in the package?

Each deployment starts with one job—billing, a claim or an IT request. The agent receives only the necessary knowledge and system access; the enterprise defines actions, approvals and human handoff. Presence combines SOPs, scoped permissions, approved actions, simulations, graders, guardrails and escalation rules.

That changes the unit of purchase. The model sits inside an operating lifecycle spanning integration, production measurement and controlled updates. Risk rises sharply when an agent moves from answering a question to changing an account or authorizing value.

  • One measurable workflow.
  • Least privilege for every tool.
  • Edge-case tests and intervention thresholds.
  • Versioned changes with human approval.

3. Reading 75% without turning it into a promise

OpenAI says Presence runs its English phone-support channel and resolves 75% of inbound issues without human help. It also says a Codex-powered improvement loop reduced handoffs by 15 percentage points in ten days. That is useful operational evidence, but vendor-reported evidence: the announcement does not publish sample size, case mix or an independent audit.

The named customers are at different stages. BBVA is exploring voice support in Mexico; SoftBank is testing natural Japanese conversations; IAG describes an opportunity. Exploration and testing are not completed production deployments.

  • Operational in the announcement: OpenAI's English support line.
  • Vendor-reported: 75% containment and a 15-point handoff reduction.
  • Not completed: BBVA exploration and SoftBank testing.
  • Not disclosed: independent results, public pricing and audited customer ROI.

4. Continuous improvement is change governance

Production sessions, escalations and quality signals expose gaps; Codex proposes changes for teams to test and approve before controlled rollout. That is stronger than ad hoc prompt editing, but it raises ownership: who may change a service policy that affects money, complaints or customer rights?

Use NIST's lifecycle: govern decision ownership, map affected people and harms, measure outcomes and edge cases, then manage rollout and rollback. Production versions need freezes, audit trails and regression suites—not only average quality scores.

5. The Saudi procurement questions

OpenAI's general enterprise commitment says business data is not used for training by default. That does not settle a Presence deployment. Contracts should specify processing location, voice and transcript retention, subprocessors, support access, deletion, incidents and transfers outside Saudi Arabia.

Saudi PDPL covers information that identifies a person, and processing includes collection, recording, storage and use. Map the call from its first second through every system action. Do not assume general API terms apply unchanged to a managed product; require a project-specific data-processing and architecture schedule.

  • Test Saudi dialects, names and addresses.
  • Disclose automation and offer immediate human routing.
  • Gate financial and sensitive actions behind appropriate verification.
  • Complete impact and transfer assessments before production.

6. A scorecard for the real economics

Do not compare only seat cost. Include discovery, integration, model and voice usage, monitoring, graders and retained human operations. Compare cost per correctly resolved case after repeat contacts, errors and remediation. With no public price, the business case must use a written quote—not an assumed API tariff.

Start with a low-risk queue and an unambiguous outcome, retain a control sample, and publish first-contact resolution, satisfaction, false handoff, unauthorized action, latency and cost weekly. Expand only after Arabic and English reach defined parity and rollback works in a live exercise.