The Sandbox Developer

Stateful vs. Stateless API Mocks in Integration Tests

Stateful mocks track state across API calls; stateless ones return canned responses regardless.

Contributing Editor · · 10 min read
Cover illustration for “Stateful vs. Stateless API Mocks in Integration Tests”
Mocking vs. Stubbing · October 8, 2026 · 10 min read · 2,167 words

Integration tests exist to check whether a sequence of API calls produces the same cumulative result that production will produce: what one call creates, a later call has to be able to find, update, and delete, in that order, with the state carried forward at every step. A unit test asks something much narrower. It asks whether a function handles a given response correctly, and a single canned response answers that question completely. Integration tests ask whether the whole system behaves correctly across a chain of calls where each step depends on what the step before it did, and no tool with no memory of prior calls can answer that question, no matter how well it answers the unit-test one. Redocly's sandbox guide calls integration "the most dangerous phase of the API lifecycle," the point where documentation meets reality, and that's why the tool used to test it has to meet reality too. The rest of this piece maps, point by point, where a popular class of testing tool fails that requirement, and what a tool has to do instead to pass it.

What a stateless mock is, and what it structurally cannot do

A stateless mock is a configuration file wearing an HTTP interface. It maps request patterns to canned responses, returns the same response every time it sees a matching pattern, and keeps no record of anything that happened on a previous call. That design is appropriate for a lot of jobs: deterministic output, no side effects, nothing to set up and nothing to tear down afterward. That means the mock cannot tell "this resource was created five calls ago" from "this resource never existed," because it never recorded either event. A POST to create a resource returns 201 because the mock's route is configured to return 201, not because anything was stored anywhere. A GET sent right after returns whatever example response that route was configured with, not the resource that was supposedly just created, because no resource was created. A list endpoint returns a fixed array that will not include the new item, because the mock never knew the POST happened. Redocly's sandbox guide states the consequence: POST a user to a mock server, and a subsequent GET will not return that user unless someone specifically scripted that exact exchange, and most teams have not scripted every permutation a real test suite needs. Integration tests must validate that sequential calls produce cumulative, consistent state, and a configuration-file mock has no mechanism for meeting that demand. Platforms that maintain state across calls and are checked continuously against the live API, Rystic among them, exist to answer the wider question a mock cannot: does the system behave correctly across a chain of calls that depend on each other?

The specific failure modes that appear only when state is missing

Mock-based integration tests can pass for months and then fail the moment they meet production, and the failure has nothing to do with a bug in the test. It happens because a passing mock test only proves that the mock's own configured responses agree with each other. That's a property of the configuration file, not of the API it's standing in for. The real API, meanwhile, might enforce a rule the mock never represented, change state in a way the mock never tracked, or show a newly created resource in a list endpoint the mock never bothered to update.

The sharpest way to see this failure is a simple sequence: POST a resource, GET it back, mutate it, then GET it again. If that final GET comes back with the state the resource had before the mutation, the tool being tested is a mock, not a stateful simulator, whatever its documentation calls it. On a genuinely stateful sandbox, the final GET returns the mutated state, because the mutation was actually stored somewhere and the read call looked it up. On a mock, the final GET returns whatever that route happens to be configured to return, usually the original example response, because there was never a link between the mutation call and the read call to begin with.

A related failure sits in webhooks. Mocks do not fire events when state changes, so a stateless mock cannot test any workflow built around a webhook that fires after a state transition, a payment getting captured, a subscription going active, a job finishing. There's nothing to configure, because the mock never tracked the transition that should have triggered the event.

A second failure sits in the state machine itself. Real APIs enforce which transitions are valid: a cancelled subscription can't be captured, a refunded payment can't be refunded twice. A mock has no model of the state machine behind its routes, so it returns whatever a given route is configured to return regardless of what came before. Invalid transitions pass silently in the test suite and then fail, visibly, in production.

API Drift and the False Confidence of a Passing Mock Suite

Hand-built mocks have to be hand-maintained: every change to the real API has to be manually copied into the stub, and most teams find out about the gap when production breaks. Mocks generated from an OpenAPI spec drift more slowly, but they're only as good as the spec feeding them. A stale or incomplete spec produces a stale or incomplete mock, and the real API quietly moves away from both without either one registering the change.

Two real cases show what that drift looks like in practice. Stripe's API version 2026-05-27.dahlia changed the billed_until field on SubscriptionItem from a field included by default to one that only appears if the caller explicitly expands it. Code that reads that field directly, reconciliation logic, renewal displays, dunning workflows, gets an empty value after the upgrade, usually without throwing any error. A mock configured before that change keeps returning the field on every call, so the test suite keeps passing while production quietly returns nothing. PayPal's IPN migration moved the webhook payload format from form-encoded to JSON. A parser still expecting the old format receives JSON instead, every field it tries to read comes back undefined, and the handler still returns HTTP 200 regardless, so PayPal marks the delivery successful even though nothing was actually processed, and the order sits there marked unpaid. A mock built against the old format has no way to surface that failure, because it was never testing the real payload shape.

Redocly's sandbox guide names the cost: once a sandbox drifts from reality, developers stop trusting it, and code that passes against the sandbox but fails in production gets blamed on the platform. Drift is a structural property of any tool that isn't checked, continuously, against the live API it claims to represent.

What Stateful Simulation Must Do for Integration Tests

A tool earns the label "stateful simulator" by meeting a specific set of behavioral requirements, not by being marketed as more capable than a mock. It has to store resources when they're created and return them on retrieval: POST creates something real, GET finds it, a list endpoint includes it, PATCH updates it, DELETE removes it, and every one of those operations changes what the next call sees. It has to enforce valid state transitions the same way the real API does, so a resource sitting in state A cannot jump to state C without passing through state B, matching whatever business logic the production system actually runs. It has to fire webhook events when state changes, because the async layer isn't optional for any API that depends on it. And it has to isolate state per test session, so one test's data never bleeds into another's, which is the only thing that makes destructive operations, deletes, refunds, cancellations, safe to run inside a CI pipeline.

Redocly's sandbox guide lists state transitions, transactional logic, and asynchronous processing as the requirements a full sandbox has to meet, and treats environmental isolation, logical or physical, as non-negotiable, since test data that leaks into production analytics or triggers a real-world action defeats the point of testing. Pre-seeded, realistic data lets a test suite immediately exercise retrieval, updates, and edge cases without a pile of setup code in front of every test, where an empty sandbox only exercises the happy path of creating something from nothing.

None of this holds up without continuous verification against the live API. A simulator that isn't regularly checked against the real API's current behavior will drift exactly like a hand-maintained mock does, producing the same false confidence, just on a longer timeline. Redocly's guide recommends treating a sandbox as a production-grade product in its own right, run through automated contract-test pipelines that deploy against the sandbox first and only promote to production once behavior matches the spec.

Where stateless mocks break in practice across specific APIs

The pattern above repeats across very different kinds of APIs. Payment APIs are the clearest example. Stripe and PayPal both require stateful testing because the payment lifecycle is a sequence of transitions, not a single request-response pair: a charge moves from pending to captured or failed, a subscription moves from trialing to active to past_due, a refund changes the charge's amount_refunded field. A mock returns whatever its route was configured for and has no way to reflect a capture that just happened a call earlier. Stripe also retries undelivered webhook events for up to three days. A serious test suite has to cover endpoint outages, replay handling, and reconciliation across that entire window, none of which a stateless mock can represent. PayPal treats a webhook as successfully delivered on any HTTP 2xx response and will retry delivery multiple times over three days; an integration that mis-parses the payload, as in the IPN migration case, will acknowledge receipt of events it never actually processed, and a mock will never catch that, because it never sees the real payload shape to begin with.

Slack bot testing runs into a different version of the same problem. Every iteration against Slack's API requires a live call, every bug needs a real workspace to reproduce, and every developer needs a test bot of their own. The Bolt SDK has had an open GitHub issue, #638, requesting proper testing support since September 2020, with a proposed solution floated in 2024 that still hadn't shipped as of April 2026. Slack also retired legacy test token creation, so teams have to migrate to apps with specifically scoped permissions, a migration a stateless mock has no way to verify. A stateful simulator of the Slack API has to track channel membership, message history, and reaction state across calls, none of which a stateless mock keeps any record of.

Kalshi's setup illustrates a related but separate problem: even a first-party sandbox, with its demo environment at external-api.demo.kalshi.co/trade-api/v2 sitting apart from production at external-api.kalshi.com/trade-api/v2, still requires network access, credentials, and tolerance for live-environment constraints that have no business inside a CI pipeline running on every commit. Stateless mocks sidestep that problem by avoiding the network entirely, but they do it by giving up the state tracking that makes the test meaningful. Platforms built specifically to simulate these kinds of APIs, Rystic included, persist what gets created through one call so a later call retrieves the actual resource that was created, and a list endpoint includes it without anyone having to script every possible permutation by hand.

Fault injection and agent testing: two failure categories that stateless mocks cannot cover at all

Fault injection, simulating latency, rate limits, 429s, 503s, and malformed responses, only means something in a stateful context. Injecting a 429 after the third call in a sequence tests a different code path than injecting it on the first call, and a stateless mock has no way to count calls or condition a fault on anything that happened earlier, because it isn't tracking anything that happened earlier. API7.ai's mocking guide draws this line explicitly, pointing to behavioral mocks, which simulate latency, rate limits, and stateful interactions, as the right tool for resilience testing, and setting them apart from static and contract-based mocks, which aren't built for that job. For payment integrations specifically, fault injection has to cover the full async failure envelope, delayed webhook events, duplicate deliveries, events arriving out of order, not just synchronous error codes. Tools that support fault injection alongside genuine state and side effects, Rystic among them, are built to catch these gaps during development, rather than letting them stay hidden until a production incident forces the question.

AI agents calling external APIs through tool use introduce a failure category that stateless mocks are particularly poorly suited to catch. An agent can generate perfectly coherent, confident text while sending malformed parameters to the API behind the scenes, and the failure stays silent: the text reads fine, the API call is wrong, and nothing in a stateless mock's design gives it any way to notice the mismatch, since it was never tracking what a sequence of calls was supposed to add up to.

More in Mocking vs. Stubbing