The Sandbox Developer

API Sandbox Environments vs. Local Simulators for CI

Local simulators solve CI repeatability problems that provider sandboxes cannot.

Contributing Editor · · 10 min read
Cover illustration for “API Sandbox Environments vs. Local Simulators for CI”
Mocking vs. Stubbing · October 10, 2026 · 10 min read · 2,187 words

Provider-hosted API sandboxes and local simulators look like substitutes for one another, but they answer different engineering questions, and treating them as interchangeable is what drags CI problems into pipelines that never needed to have them. A sandbox built by Stripe, PayPal, or Kalshi is a stateful environment that runs on the provider's own infrastructure, built so a developer can experiment with real API behavior without touching production data or production money. Its job is to isolate consequences: a test charge doesn't bill a card, a test message doesn't reach a real customer. A CI pipeline exists for the opposite reason. Its job is to isolate circumstances, so that every commit or pull request spins up a disposable environment that mirrors production closely enough to be meaningful, yet produces the identical result on every run regardless of what else is happening on the network or in the provider's infrastructure that day.

The qualities that make a provider sandbox excellent for onboarding, namely a real network, a shared persistent world, and state the provider manages, are the exact qualities that turn it into friction once it's dropped into a CI pipeline. A new engineer exploring the Stripe API for the first time benefits from a shared, persistent, realistic environment. It needs an environment that behaves the same way every single time, with no dependency on who else is using the sandbox at that moment or what state they left it in. That mismatch, not any defect in how well a provider builds its sandbox, is the starting point for everything that follows.

Four structural properties of provider sandboxes that break CI repeatability

The gap between sandboxes and CI isn't a quality problem that a better-funded provider could engineer away: it's a structural incompatibility that shows up in four distinct ways.

Network dependency is the first. Interacting directly with a live, provider-hosted API during development carries real risk: unintended side effects, data corruption, unexpected charges. In a CI context, that risk turns into something worse than danger. It turns into ambiguity. A flaky network connection or a provider-side outage produces a flaky build, and there is no clean way to tell from the failed run whether the team's own code broke or the provider's infrastructure simply hiccupped.

Rate limits are the second. A team running dozens of parallel pull-request builds against a shared sandbox tier can exhaust that allowed concurrency well before the day's work is done, and the pipeline queues up behind a limit that has nothing to do with the team's own code.

Shared state is the third, and it's the one that produces the most confusing failures. Most flaky API test suites are flaky precisely because they are unintentionally stateful: tests leak state into one another across runs, and nobody wrote that behavior on purpose.

Non-determinism is the fourth. None of this is a defect; it reflects real-world delivery guarantees a production system has to tolerate. But timing-dependent behavior inside a shared environment is precisely what a reproducible test suite cannot absorb. A test that passes because a webhook arrived in the expected order on Tuesday can fail on Wednesday for no reason connected to the code under test.

Stateful simulation and the cost of stateless integration tests

The requirement CI-grade integration testing actually has isn't simply speed or independence from the network. The environment has to remember state across calls the same way a real API does. That consistency, when it appears in a stateless mock, is a property of how carefully someone configured the mock for that particular test. It is not a property of the API being simulated, and it will not generalize to the next test someone writes.

A stateful environment behaves differently because it maintains a persistent world across the sequence of calls a test makes. A mock that always returns a hardcoded subscription object will pass every test in the suite even if the subscription creation call was never actually made. That produces false confidence at the integration layer, which teams rely on most to catch the bugs that unit tests miss.

How the two-track pipeline pattern handles this in practice

Teams that integrate external APIs already run something close to a two-track pattern, often without naming it as a deliberate architectural choice. The common shape puts lightweight, stateless stubs in the early phases of CI, run on every pull request, and shifts to more integrated, higher-fidelity checks in later-stage environments that run less frequently.

The decision to run the real provider sandbox nightly instead of on every PR is a cost and latency decision, not a preference. But the network dependency and the shared-state risk that come bundled with a live sandbox make it a poor fit for something that needs to run on every single commit. Running it nightly confines the blast radius of a flaky build to a single scheduled run.

That compromise leaves a gap. What runs on every PR, the stateless stub, is too shallow to catch the class of bugs that live in state transitions. What runs nightly, the live provider sandbox, is too slow and too fragile to run at PR frequency. A stateful local simulator is the piece missing from the per-PR slot: something with the behavioral depth of the sandbox and the speed and isolation of the stub.

What locally-run stateful simulators are built to do differently

A locally-run stateful simulator addresses the four structural CI problems by design, rather than papering over them with retries and workaround scripts.

It carries no rate limits in the way a provider sandbox does, because it isn't a shared resource. Determinism follows the same logic. The simulator returns what the test expects because the simulator's behavior is under the test's control, with no three-day message retention window to worry about and no risk of webhooks arriving out of order.

Pre-seeded, realistic data matters more than it might first appear. A simulator that starts from a cohesive, realistic world removes that extra layer and lets the test focus on the behavior it's actually meant to check.

Rystic's simulator approach compared to other local options

Among the available approaches to local simulation, the question that separates a durable solution from a liability is whether the simulator is continuously checked against the live API it's standing in for. Without that check, any local stand-in will drift quietly away from reality and generate false confidence that is worse than an outright test failure.

Rystic builds stateful simulators for Slack, GitHub, Stripe, Kalshi, PayPal, Resend, Lob, Okta, and Linear, all locally hosted and runnable without live credentials. Each simulator ships with pre-seeded, realistic data, so a test suite begins from a meaningful world. Rystic publishes fidelity scores across its catalog, covering APIs including Slack, GitHub, Kalshi, PayPal, and Lob, giving a team something it can audit to verify the claim. The simulators scale horizontally across CI containers without shared-state contention between parallel jobs, since each instance is its own isolated world.

Hand-written mocks are fast and cheap to produce for a single endpoint; they carry no external dependency. Left unattended, they drift silently as the real API they imitate continues to evolve, since nothing in their construction keeps them honest against that evolving target. They remain well suited to isolating a single call at the unit level, but they don't provide integration-layer coverage on their own.

Record-and-replay tools capture real API responses once and play them back later. These tools also can't simulate faults or inject error conditions without manual editing of the captured data. They work reasonably well for read-heavy, stable endpoints, and they become fragile quickly for anything write-heavy or still evolving.

The provider sandbox, included here for completeness, offers the highest fidelity of any option, since it is quite literally the provider's own infrastructure running the real code paths. It remains the right tool for manual end-to-end verification, onboarding, and pre-launch smoke testing. It also carries, undiminished, all four structural CI problems described earlier: network dependency, rate limits, shared state, and non-determinism. It belongs in a nightly or pre-release slot, not wired into every pull request.

Fidelity drift as the silent failure mode for any non-verified simulator

A local simulator that isn't checked continuously against the live API it imitates will eventually tell the developer something false, and because that falsehood produces a passing test rather than a failing one, it does more damage than an honest failure would. The most common failure mode for hand-maintained mocks and static simulators is staleness: the real API adds a required field, changes an error code, or reorders part of a response, and the mock keeps returning yesterday's answer indefinitely, with nothing in its design to flag the divergence.

The consequence lands at the worst possible time. A developer sees a green CI run, ships the change with confidence, and the code meets the real API for the first time only once it's already in production. The drift that had been accumulating quietly inside the mock becomes visible in production, exactly when it's most expensive to discover.

Continuous verification changes that picture by running probes against the real API before each simulator release and publishing the results of those probes. Trusting that a mock someone wrote six months ago still reflects the current state of the API is an assumption, and production is where it eventually gets tested.

Fault injection as a CI requirement, not an optional extra

A test suite that only checks the happy path against a well-behaved simulator leaves out the conditions that actually break real launches. Those conditions are asynchronous, partial, and dependent on retry behavior, and a simulator earns its place in CI by producing them on demand. A test plan that never exercises an intermittent endpoint failure, or a workflow that's interrupted partway through and has to resume safely, is a demonstration that the system works when nothing goes wrong, which is a narrower claim than most teams believe they're making.

The categories of fault that belong inside CI, not reserved for a manual pre-launch exercise, include network throttling, rate-limit responses in the form of 429s, service-unavailable 503s, duplicate webhook delivery, out-of-order event delivery, invalid token responses, and oversized payloads. A simulator built with fault injection as a first-class feature turns resilience testing into something routine, run in the same CI pass as the happy-path tests, rather than a manual exercise that only happens once, right before a major release, when there's the least time to fix what it finds.

Applying this to AI agents and RL environments that call external APIs

AI agents that call external APIs inherit every CI requirement that traditional software already has, and add new ones on top that make the shared-sandbox model even harder to justify. An LLM agent working through a multi-turn workflow against Slack, GitHub, or Stripe has to see state persist correctly across its own tool calls: something written in turn one must be visible to a read in turn three. A stateless mock, returning the same hardcoded response regardless of what the agent did a moment earlier, cannot validate that behavior.

Evaluating an agent at CI scale typically means running a large number of parallel episodes against the simulated environment. Reinforcement learning environments sharpen the same requirement further: they need deterministic, controllable state that resets cleanly between episodes and scales across potentially thousands of containers running at once. Those are the identical properties that make a stateful local simulator suitable for ordinary CI. The same architecture turns out to be the correct substrate for RL training as well, not a coincidental overlap.

Provider sandboxes carry a specific unsolved problem for agent testing. An agent can write code and run its own unit tests against such a sandbox, but genuine end-to-end verification remains a manual step performed by a human, unless the agent is connected to an environment that's both governed and real, which a locked-down provider sandbox generally isn't built to be.

Keeping or replacing the provider sandbox

The right move isn't to remove provider sandboxes from the workflow. It remains the reference point a team checks a simulator against, confirming that the simulator's behavior still matches the real API.

A local stateful simulator takes over everywhere repeatability and scale matter most: every pull request and every commit, where network dependency and shared state would otherwise kill reproducibility; parallel CI runs, where rate limits and concurrency ceilings would throttle the pipeline rather than the code under test; fault injection and resilience testing, where deterministic, programmable error conditions have to be available on demand; AI agent and RL evaluation, where per-episode state reset and horizontal scale aren't conveniences but requirements; and offline or air-gapped environments, where network egress isn't an option.

The two-track pattern many teams already practice informally, with a stateful local simulator running on every PR and a provider sandbox reserved for a nightly or pre-release slot, is the right model precisely because each environment does the job it was built for. Neither one is a substitute for the other. Each is a tool suited to a particular question, and the pipeline works once the right question goes to the right tool.

More in Mocking vs. Stubbing