OpenWeatherMap
openweathermap.org · Developer / Data APIs modeIntegrated
Last probed 2026-08-16 17:31 UTC · Provisional
OpenWeatherMap — Self-serve weather data API, free-tier key.
Verdict: An agent can activate and manage on its own. Discover and operate aren't assessed yet.
- DiscoverNot assessed
- ActivateEasy
- OperateNot assessed
- ManageEasy
Capability parity
Each stage answers one question about what an agent can do unattended. Stages are ordered by the credential-gateway dependency: Activate issues the keys that unlock Operate and Manage.
Discover
Can an agent find and read it, no account?
- Be found by a retrieval agentNot assessed
- Read what it does and how to use itNot assessed
- Read pricing & limits without loginNot assessed
Activate
Can an agent sign up and walk away with the keys it needs?
- Create an account self-serveEasy
Signup inspected — no blocker found.
Evidence
- Path tried
- signup flow
- Where it stalled
- completed
- Evidence tier
- Tier 1 (unauth probe)
- Verification
- Evidence-only — graded from passive signal, not live-verified.
- Obtain an API / MCP access credentialEasy
Self-serve credential at signup.
Evidence
- Path tried
- signup → key issuance
- Where it stalled
- completed
- Evidence tier
- Tier 1 (unauth probe)
- Verification
- Evidence-only — graded from passive signal, not live-verified.
- One-time account / payment-method provisioningSignupOne-time account + API key signup required before any call succeeds; no CAPTCHA, no phone/email OTP, no manual review observed.
- Obtain a management / admin keyNot assessed
- Accept the terms without a bot-banEasy
No bot-prohibition signal found.
Evidence
- Path tried
- ToS / policy scan
- Where it stalled
- completed
- Evidence tier
- Tier 0 (passive scan)
- Verification
- Evidence-only — graded from passive signal, not live-verified.
Operate
Can an agent do the service's actual jobs — call, transact, pay?
- Call the primary service over MCP / APINot assessed
- Complete a transaction / checkoutNot applicable
Integrated tool — the API call is the transaction, no separate checkout.
- Pay via an agent rail (x402 / ACP / AP2)Not assessed
Manage
Can an agent administer the account and provision resources?
- Provision / scale / tear down a serviceEasy
Self-serve provisioning signal at signup.
Evidence
- Path tried
- docs → provisioning API
- Where it stalled
- completed
- Evidence tier
- Tier 1 (unauth probe)
- Verification
- Evidence-only — graded from passive signal, not live-verified.
- Change account details / set budget / rotate keysNot assessed
- View usage & billingNot assessed
- Cancel / export / delete (offboard)Not assessed
Findings
How scoring works →Interrupt tier: Low-Interrupt — derived from 1 finding; interrupt score 80 / 100 — the badge-driving metric. (range 48.4–98.4 · 7 unassessed)
Every tool starts at 100. Each interrupt found subtracts a penalty of base weight × timing × hardness × evidence — the multiplier values below come from the same rubric tables the score is computed with. New snapshots also provisionally subtract 50% of the full worst-case tax for each unassessed property and 25% for evidence-graded successes not live-verified by Tier 2.
Property coverage
A check only counts as clean when a probe tier capable of detecting that interrupt actually ran and found nothing. Not tested ≠ clean.
| Property | Stage | State | Evidence | Required |
|---|---|---|---|---|
| One-time account / payment-method provisioning | manage | failed | Active finding — itemized in the penalty matrix below. Tier 2 assesses this property. | no |
| Mandatory phone verification call | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| Mailed physical code (postal) | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| “Contact sales” wall (no self-serve path) | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| Manual approval queue | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| Rate-limit-triggered manual review | manage | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| Identity / KYC verification | activate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| Crypto wallet / key-custody setup | operate | unassessed | No probe tier that can detect this has run yet — unknown, not clean. Tier 2 assesses this property. | no |
| CAPTCHA / bot-challenge | activate | success | None found — checked at Tier 0 (passive scan) + Tier 1 (unauth probe) (evidence-graded; not yet live-verified by a Tier 2 probe) Tier 0 assesses this property. | no |
| Email OTP verification | activate | success | None found — checked at Tier 1 (unauth probe) (evidence-graded; not yet live-verified by a Tier 2 probe) Tier 1 assesses this property. | no |
| SMS / phone OTP verification | activate | success | None found — checked at Tier 1 (unauth probe) (evidence-graded; not yet live-verified by a Tier 2 probe) Tier 1 assesses this property. | no |
| ToS clause banning automated/bot access | discover | success | None found — checked at Tier 0 (passive scan) (evidence-graded; not yet live-verified by a Tier 2 probe) Tier 0 assesses this property. | no |
| Mandatory redirect to vendor’s own site | operate | success | None found — checked at Tier 1 (unauth probe) (evidence-graded; not yet live-verified by a Tier 2 probe) Tier 1 assesses this property. | no |
| SSO-only signup, no API-key issuance | activate | success | None found — checked at Tier 1 (unauth probe) (evidence-graded; not yet live-verified by a Tier 2 probe) Tier 1 assesses this property. | no |
| No Machine Interface | operate | success | None found — checked at Tier 1 (unauth probe) (evidence-graded; not yet live-verified by a Tier 2 probe) Tier 1 assesses this property. | yes |
| Interrupt found | Base weight | × Timing | × Hardness | × Evidence | = Penalty |
|---|---|---|---|---|---|
| Starting score | 100 | ||||
| One-time account / payment-method provisioningSignupOne-time account + API key signup required before any call succeeds; no CAPTCHA, no phone/email OTP, no manual review observed. | 2 | ×1one-time setup | ×1soft | ×0.8Tier 1 (unauth probe) | −1.6 |
| Total penalty | −1.6 | ||||
| Full unknown tax 93 — capped at 25 (point score uses 50% of the capped value) | −93 | ||||
| Evidence-graded tax (successes not live-verified by a Tier 2 probe) 161.4 — capped at 25 (point score uses 25% of the capped value) | −161.4 | ||||
| Raw score = 100 − 1.6 = 98.4 Point score = 98.4 − 50% × 25 − 25% × 25 = 79.7 → 80 | 80 | ||||
How we scored this
- Confidence: Provisional (0.55) — Based on Tier 0/1 evidence only (passive checks, no live transaction attempt) — treat the number as a triage signal, not a verified outcome.
- Machine-Readability sub-score: 88 / 100 (a separate integration-ease signal — it never enters the Interrupt Score or badge).
- Some capability stages are not yet assessed — shown plainly in the spine above, never as a pass.
- Rubric, tiers, confidence model, and badge thresholds →
Machine-readability checks not yet recorded
Four check groups, equal weight: machine interfaces, browsing efficiency, element discoverability, and discovery/agent access. Each check carries an evidence grade (proven / plausible / emerging) — an honesty label that never weights the score. What these mean →
No per-check breakdown recorded for this listing.
Evidence & history
Last probed 2026-08-16 17:31 UTC. The raw JSON is the archival record.
This is the first assessment on record.
| Date | Score | Badge | Rubric version |
|---|---|---|---|
| 2026-08-16 17:31 UTC | 80 | Low-Interrupt (current) | [email protected] |
Alternatives in Developer / Data APIs
Other Developer / Data APIs tools in the same mode (integrated), ranked by Overall LLM-Friendliness (the mean of Interrupt Score and Machine-Readability) — comparable to this listing because they share its usage mode.