Built on 12 years · 150M requests/month
Real APIs are dangerous, rate-limited, and never the same twice. ReqRes is the deterministic world where your agent practices: seeded data, deliberate failures, byte-identical replays. The safe place to break things before production.
curl-able in 10 seconds 15 scripted failure modes known to every frontier model
An agent practicing retries against a real API is an incident waiting for a timestamp. Rate limits, real money, real customer data. None of it belongs in a training loop.
Hand-rolled fixtures drift from reality the day you write them. ReqRes serves living, relational, versioned data (users, orders, auth), maintained so you never have to.
If the same request can return different bytes, your eval can't tell a model regression from network weather. Determinism makes agent failures reproducible, and therefore fixable.
Cursor pagination, ULIDs, 20+ fields nested three levels deep, nullable on purpose. 247 users per seed, infinite seeds.
Relational data that references users, products, and addresses, including line items pointing at deleted products, because real data does that.
Returns a session or an MFA challenge depending on the email. Shape-varying responses, with hypermedia hints for the next step.
Fifteen deliberate failure modes on demand: the catalogue below. Every one tagged X-Agent-Sandbox-Intentional so your error reporter stays quiet.
A replayable record of every call your agent made: exportable, diffable, attachable to an eval run. For eval teams, the log is the product.
Point any model at /llms.txt, /llm.txt, or /openapi.json, or hand it the MCP server. No human docs required.
Chain them: scripted scenario sequences: succeed twice, then 429, then timeout. Retry logic that survives contact with reality.
In build · Q4 2026
A Stripe-shaped API where agents practice the scariest integration in software, without touching money.
Queued · 2027
Slack-shaped events, rate limits, and retry semantics for agents that live in channels.
Queued · 2027
Salesforce-shaped objects, auth scopes, and query semantics for enterprise agent workflows.
Simulated interfaces for practice and evaluation. Not affiliated with or endorsed by the services they resemble.
Every quarter we score the major coding agents against the full failure catalogue (pagination, retries, malformed JSON, auth flows) using trajectory logs from the sandbox. Published in full, methodology open.
| Agent | Pagination | Retry & backoff | Error handling | Overall |
|---|---|---|---|---|
| Claude Code | — | — | — | — |
| Cursor | — | — | — | — |
| Copilot | — | — | — | — |
| Devin | — | — | — | — |
First results: September 2026. Get notified when the leaderboard goes live →
Free
$0
Kick the tyres. No card, no signup.
Agent Developer
$49/mo
For teams shipping agents that hit APIs.
Agent Platform
Custom
For coding-tool companies, eval harnesses, and labs.
We're onboarding a small group of environment builders, eval teams, and agent companies before public launch: free Agent Platform access through Q1 2027, a direct line on the roadmap, and your logo on this page when we go loud.
5 of 10 slots remaining