ChatGPT Instant Checkout / OpenAI ACP
openai.com · Agent-Commerce Infrastructure modeIntegrated
Last probed 2026-09-08 04:41 UTC · Verified (Documented)
ChatGPT Instant Checkout / OpenAI ACP — In-chat shopping that hands off to the merchant's own checkout to complete purchase.
Verdict: Activate is only partly probed. Discover, operate, and manage aren't probed yet.
- DiscoverNot yet probed
- ActivateIncomplete
- OperateNot yet probed
- ManageNot yet probed
Capability parity
Each stage answers one question about what an agent can do unattended. Stages are ordered by the credential-gateway dependency: Activate issues the keys that unlock Operate and Manage. A row marked “Not yet probed” means no probe has tried that capability yet — a different ledger from the “unassessed” properties under Findings, where no probe tier capable of catching that interrupt has run.
Discover
Can an agent find and read it, no account?
No probe has attempted this stage yet: Be found by a retrieval agent; Read what it does and how to use it; Read pricing & limits without login.
- Be found by a retrieval agentNot yet probed
Retrieval access not confirmed - no robots/sitemap/discovery check passed.
- Read what it does and how to use itNot yet probed
No machine-readable docs found - llms.txt, AGENTS.md, and schema.org checks all failed or are absent.
- Read pricing & limits without loginNot yet probed
Public pricing was not confirmed - the page was absent, gated, or unreadable.
Activate
Can an agent sign up and walk away with the keys it needs?
Assessment incomplete - not yet probed: Create an account self-serve; Obtain an API / MCP access credential; Obtain a management / admin key.
- Create an account self-serveNot yet probed
Signup flow not yet inspected (Tier 1 has not run).
- Obtain an API / MCP access credentialNot yet probed
No self-serve credential signal detected - key issuance not yet probed.
- Obtain a management / admin keyNot yet probed
No scoped or administrative key path was confirmed after activation.
- Accept the terms without a bot-banEasy
No bot-prohibition signal found.
Evidence
- Path tried
- ToS / policy scan
- Where it stalled
- completed
- Evidence tier
- Tier 0 (passive scan)
- Verification
- Evidence-only — graded from passive signal, not live-verified.
Operate
Can an agent do the service's actual jobs — call, transact, pay?
No probe has attempted this stage yet: Call the primary service over MCP / API; Pay via an agent rail (x402 / ACP / AP2).
- Call the primary service over MCP / APINot yet probed
No machine interface found by passive checks, and no live Tier-2 call attempted.
- Complete a transaction / checkoutNot applicable
Integrated tool — the API call is the transaction, no separate checkout.
- Pay via an agent rail (x402 / ACP / AP2)Not yet probed
No agent-payment metadata detected - payment rails not yet probed.
Manage
Can an agent administer the account and provision resources?
No probe has attempted this stage yet: Provision / scale / tear down a service; Change account details / set budget / rotate keys; View usage & billing; Cancel / export / delete (offboard).
- Provision / scale / tear down a serviceNot yet probed
Not probed - provisioning is not part of the probe suite yet.
- Change account details / set budget / rotate keysNot yet probed
Not probed - account admin is not part of the probe suite yet.
- View usage & billingNot yet probed
Authenticated usage and billing access was not confirmed.
- Cancel / export / delete (offboard)Not yet probed
No self-serve cancellation, export, or deletion path was confirmed.
Findings
How scoring works →Interrupt tier: Moderate-Interrupt — derived from 1 finding; interrupt score 60 / 100 — the badge-driving metric. (range 47.8–72.8 · 14 unassessed)
Every tool starts at 100. Each interrupt found subtracts a penalty of base weight × timing × hardness × evidence — the multiplier values below come from the same rubric tables the score is computed with. New snapshots also provisionally subtract 50% of the full worst-case tax for each unassessed property and 25% for evidence-graded successes not live-verified by Tier 2.
Property coverage
A check only counts as clean when a probe tier capable of detecting that interrupt actually ran and found nothing. Not tested ≠ clean.
| Property | Stage | State | Evidence | Required |
|---|---|---|---|---|
| Mandatory redirect to vendor’s own site | operate | failed | Active finding — itemized in the penalty matrix below. Tier 1 assesses this property. | no |
| CAPTCHA / bot-challenge | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 0 assesses this property. | no |
| Email OTP verification | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 1 assesses this property. | no |
| SMS / phone OTP verification | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 1 assesses this property. | no |
| Mandatory phone verification call | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| Mailed physical code (postal) | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| “Contact sales” wall (no self-serve path) | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| Manual approval queue | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| ToS clause banning automated/bot access | discover | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 0 assesses this property. | no |
| SSO-only signup, no API-key issuance | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 1 assesses this property. | no |
| Rate-limit-triggered manual review | manage | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| One-time account / payment-method provisioning | manage | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| Identity / KYC verification | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| Crypto wallet / key-custody setup | operate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| No Machine Interface | operate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 1 assesses this property. | yes |
| Interrupt found | Base weight | × Timing | × Hardness | × Evidence | = Penalty |
|---|---|---|---|---|---|
| Starting score | 100 | ||||
| Mandatory redirect to vendor’s own siteEvery transactionEvery purchase hands off to the merchant's own checkout page to complete (publicly documented product behavior, post-rollback Instant Checkout / Agentic Commerce Protocol model). | 4 | ×8every transaction | ×1soft | ×0.8Documented | −27.2 |
| Total penalty | −27.2 | ||||
| Full unknown tax 253.2 — capped at 25 (point score uses 50% of the capped value) | −253.2 | ||||
| Raw score = 100 − 27.2 = 72.8 Point score = 72.8 − 50% × 25 − 25% × 0 = 60.3 → 60 | 60 | ||||
What would make this Easy
One concrete change per blocker found. These are the diffs between this listing's current parity and an unattended agent path.
- Operate — Mandatory redirect to vendor’s own siteExpose the transaction over an API / MCP so checkout doesn't require a browser redirect.
How we scored this
- Confidence: Verified (Documented) (0.8) — Backed by strong public-record evidence of real end-to-end behavior, without our own live probe run.
- Machine-Readability sub-score: 90 / 100 (a separate integration-ease signal — it never enters the Interrupt Score or badge).
- Some capability stages are not yet probed — shown plainly in the spine above, never as a pass.
- Rubric, tiers, confidence model, and badge thresholds →
Machine-readability checks not yet recorded
Four check groups, equal weight: machine interfaces, browsing efficiency, element discoverability, and discovery/agent access. Each check carries an evidence grade (proven / plausible / emerging) — an honesty label that never weights the score. What these mean →
No per-check breakdown recorded for this listing.
Evidence & history
Last probed 2026-09-08 04:41 UTC. The raw JSON is the archival record.
This is the first assessment on record.
| Date | Score | Badge | Rubric version |
|---|---|---|---|
| 2026-09-08 04:41 UTC | 60 | Moderate-Interrupt (current) | rubric v2026.4 |
Alternatives in Agent-Commerce Infrastructure
Other Agent-Commerce Infrastructure tools in the same mode (integrated), ranked by Overall LLM-Friendliness (the mean of Interrupt Score and Machine-Readability) — comparable to this listing because they share its usage mode.
- 95Overallx402 (Coinbase protocol)
x402.org · Agent-Commerce InfrastructureMR92 - 88OverallFewsats
fewsats.com · Agent-Commerce InfrastructureMR78 - 78triageOverallNekuda
nekuda.ai · Agent-Commerce InfrastructureMR74 - 72Overallx402 Bazaar
coinbase.com · Agent-Commerce InfrastructureMR64 - 69triageOverallCrossmint
crossmint.com · Agent-Commerce InfrastructureMR62