Testing Stripe Payment Flows With a Stateful Local Simulator
A stateful local simulator solves the consistency problems that plague Stripe integration testing.

A Stripe payment is never a single request. A charge depends on a customer, a refund depends on a charge, a subscription depends on a product and a price, and each step in that chain only works if the step before it left the right state behind. That dependency chain is the reason stateless mocks and shared sandboxes both struggle to test Stripe integrations well, and it's the starting point for understanding what a stateful local simulator actually fixes.
Stateless mocks and shared sandboxes as a poor fit for Stripe payment flows
Start with the sequence itself. A customer object has to exist before a payment method can attach to it. A PaymentIntent has to reference that payment method before it can confirm. A charge only comes into being once the PaymentIntent succeeds, and a refund only makes sense against a charge that already happened. None of this is an accident of API design. Stripe's objects reference each other by ID, and the API enforces those relationships at every step, so a test that only checks whether a single endpoint returns the right shape of response tells you nothing about whether the objects it touches actually connect to each other the way they would in production.
Stripe's own documentation points at the problem without naming it. The Billing testing guide recommends using separate sandboxes for local development and continuous integration so automated tests don't interfere with products, prices, customers, webhook endpoints, settings, or test data. That recommendation only makes sense if sandbox state is shared, mutable, and persistent across runs, which it is: a customer or subscription left behind by one test run can quietly corrupt the assertions of the next one. Splitting sandboxes between local work and CI reduces how often that happens, but it doesn't remove the deeper constraint. Every sandbox still needs an internet connection and a live Stripe account, and its API access is capped by rate limits. Stripe's testing overview confirms that sandboxes typically run on test API keys tied to an account, and while Stripe's CLI now supports provisioning anonymous sandboxes without registering an account, offline or air-gapped testing still isn't possible through this route.
Stateless mocks run into a different wall. A mock server has no memory between requests, so whatever consistency appears to exist between the first call in a test and the third call is just a reflection of how someone configured the mock, not a property the real API would enforce. If a mock returns a hardcoded customer in step one and a hardcoded subscription in step three, nothing has verified that the subscription actually references the customer created earlier. The link looks real in the test output and is a fiction. Stripe's own stripe-mock tool is explicit about this limitation: create a customer through it and the tool forgets that customer on the very next call, a PaymentIntent never progresses through its state machine, and a charge carries no real connection to the customer who supposedly made it. That's a reasonable design for checking response schemas. It's the wrong tool for proving that a payment flow holds together end to end.
The meaning of "stateful" for a Stripe simulator
A stateful simulator remembers the cumulative result of every call made during a session, and it enforces that result on every call that follows. That's a specific, checkable property, not a marketing adjective. Create a customer, and a later list-customers call has to return it. Subscribe that customer to a plan, and pulling up the customer record has to show the subscription attached to it. Cancel the subscription, and the customer record has to reflect the cancellation on the next read. None of this requires a richer config file. It requires a running system that behaves like a small, consistent world rather than a script that answers questions in isolation.
localstripe is a useful illustration of the distinction because its own documentation draws it directly: unlike stripe-mock, localstripe is stateful, tracking actions performed so that those actions affect the queries that follow, which makes it usable as an end-to-end test server. localstripe gives a concrete example worth walking through. Create a plan. Create a customer. Subscribe that customer to the plan. Then retrieve the customer record, and the subscription shows up nested inside it, exactly the way real Stripe would return it. Nothing about that result was hardcoded into the third call. It fell out of the first two.
Three specific capabilities separate a simulator like this from a stateless mock. The first is object persistence: records created in one request have to survive and show up correctly in later reads, lists, and updates. The second is enforcement of relationships and state-machine transitions: a PaymentIntent actually moves through its real sequence, from requires_payment_method to requires_confirmation to requires_action if extra authentication is needed, then to processing and on to succeeded, or back to requires_payment_method if the charge is declined, based on what the test actually sends it. The third is webhook emission: state changes have to fire signed webhook events to a registered endpoint, so an application's event-handling code can be tested without any connection to real Stripe. localstripe supports registering webhook endpoints through a special config route and signs the events it sends with a provided secret, matching the same contract real Stripe webhook delivery uses.
None of this holds up as a testing fixture without one more piece: a way to start from the same place every time. Deterministic seeding is what makes a stateful simulator reliable. Without it, a stateful simulator still carries the same contamination risk as a shared sandbox, because whatever a previous run left behind becomes the starting condition for the next one. Pairing seeding with a reset endpoint lets every test or CI job boot an identical, known world and check its assertions against exact starting conditions. localstripe exposes this as a DELETE /_config/data endpoint: flushing it before a run clears everything and restores a clean slate.
The full Stripe payment lifecycle mapped onto a stateful simulator
A production Stripe integration rarely touches just one object type. A typical flow runs through customers, payment methods, PaymentIntents, charges, and refunds at minimum, and all five have to stay correctly linked for a simulator to be worth testing against. Start with the simplest case: confirming a PaymentIntent against a customer's attached payment method creates a charge, and that charge has to carry both the PaymentIntent and the customer reference forward into every later read. If a simulator loses that linkage, every assertion downstream of it is checking a fiction.
Refunds add a layer of detail that's easy to get wrong. A refund reduces a charge's amount_refunded field and updates its refunded flag, but that flag only flips to true once the charge is fully refunded. A partial refund leaves it false. If a simulator doesn't track this distinction, you can't write a meaningful refund-idempotency test, because the test has no reliable signal to check against.
Billing is where the object graph gets deepest. A subscription attaches to a price, a price attaches to a product, and generating an invoice at the end of a billing period depends on all three staying present and consistent with each other. Billing also introduces a kind of state that's awkward to test against a live sandbox: subscriptions generate invoices on a schedule, and no CI pipeline can afford to wait for a real billing period to elapse. Stripe's Billing testing guide addresses this directly with test clocks: they simulate billing objects like subscriptions moving through time inside a sandbox, so real calendar days don't have to pass. Setting one up takes no code, since simulations can be created straight from the Dashboard, but the mechanism still depends on a connection to a remote Stripe sandbox. A local simulator that implements its own clock-advance mechanism removes that dependency: advance simulated time by whatever interval a test needs, check the invoice that gets generated, and verify its line items, all without a network call.
The Billing testing guide also lists the asynchronous events a real billing integration has to handle correctly: invoice.created, invoice.finalized, invoice.finalization_failed, invoice.paid, invoice.payment_action_required, invoice.payment_failed, invoice.upcoming, invoice.updated, customer.subscription.deleted, customer.subscription.paused, customer.subscription.resumed, and customer.subscription.updated. A simulator earns its keep for webhook testing only if it fires all of these at the correct state transitions, not a representative sample.
Idempotency sits apart from the rest of the object graph because it tests behavior, not data. A well-built integration has to prove that submitting the same PaymentIntent confirmation twice doesn't create two charges. Real Stripe enforces this through the Idempotency-Key header, so if you want to test an application's retry logic without risking a duplicate charge, you need a simulator that tracks those keys and returns the original cached response when it sees a repeat. A stateless mock simply answers every call the same way it was configured to, with no way to tell a first submission from a retry. This means the double-charge scenario, the exact failure idempotency keys exist to prevent, can't be tested.
Test isolation and CI pipeline design when the simulator is the dependency
Once the simulator runs locally, shared sandbox state no longer leaks between runs and contaminates test results. Each CI job can start its own fresh, seeded instance instead of competing with every other job for the same shared account, so an entire category of flaky test failures tied to leftover state from a previous run just goes away. Stripe's own advice to keep separate sandboxes for local development and CI is itself an acknowledgment that shared sandboxes corrupt each other's data across runs. A local simulator sidesteps the problem.
In practice, that looks like starting a fresh container per job, localstripe supports this with a single Docker command, docker run -p 8420:8420 adrienverge/localstripe:latest, and flushing it at teardown so every run begins from a guaranteed clean state. Deterministic seeding takes this further: instead of starting from an empty world, a job can start from a pre-populated one, with specific customers, payment methods, and existing subscriptions already in place, so tests assert against known IDs and known states instead of constructing every fixture inline before each run.
This capability maps naturally onto how a pipeline should be staged. Unit tests need no external dependency. Contract tests run against the stateful simulator: it's fast enough to run on every pull request, needs no network connection, and can represent the full object graph those tests need to check. End-to-end smoke tests run against a live Stripe sandbox, but only as a gate shortly before a production deploy, so they confirm the integration's wire format still matches what real Stripe returns, not a dependency for every CI run. Structured this way, rate limits and live credentials never sit on the critical path for day-to-day development or pull-request checks. They're reserved for a narrow, infrequent gate where they belong.
Isolation within a single test suite follows the same logic as isolation between CI jobs, and localstripe's DELETE /_config/data endpoint flushes all stored data on command, so a teardown hook can restore a clean state between individual test cases without restarting the simulator's process each time.
Testing the decline and error paths that define real payment resilience
Everything above assumes payments that succeed. Resilience testing is about what happens when they don't, and Stripe's decline behavior is more specific than a generic error response. When a real payment declines, the PaymentIntent goes back to requires_payment_method and carries a last_payment_error object that describes what went wrong. The associated charge records a status of failed. The HTTP response itself comes back as 402, with a card_error type (distinct from the error's code field) and a payment_intent field embedded directly in the error body, causing Stripe's SDKs to raise a typed CardError.
A stateless mock can return a 402 with a hardcoded error body well enough, but it can't update the PaymentIntent's status or write a failure record onto the charge. Any later read the application makes, polling the PaymentIntent's status after the decline, for instance, comes back inconsistent with the error the application just received. A stateful simulator that implements the full decline path, moving the PaymentIntent back to requires_payment_method, recording the failed charge, and embedding the error correctly in the response, keeps every downstream read consistent with the failure the application just handled.
Stripe's test card numbers encode specific decline scenarios: cards that trigger insufficient funds, a generic decline, or a 3D Secure challenge, among others. If a simulator recognizes these same card numbers and produces the matching state transitions, it brings that entire test-card contract offline, so you can exercise error-handling code against the exact card numbers you already use against live Stripe, with no fixture changes when you switch between the two environments.
Fault injection at the network layer is a related but separate concern that applies to the connection itself. Because a local simulator runs as a real HTTP server on localhost, you can put a network proxy such as Toxiproxy between the application and the simulator to inject latency or drop connections mid-request without touching either side. That setup makes it possible to test an application's timeout and retry logic directly: specifically, whether a timeout during a PaymentIntent confirmation causes a retry to create a second charge. Verifying that requires a simulator that tracks idempotency keys, since the test's entire premise rests on one specific hypothesis: if the HTTP connection drops after the simulator has received the confirmation request but before the application receives the response, a retry carrying the same idempotency key has to return the original result. A stateless mock has no way to hold that hypothesis up to scrutiny, because it has no memory of the first request to compare the retry against.