Contract Testing With Pact for Third-Party APIs
Pact's verification assumes both teams cooperate, breaking down when the provider is unreachable.

Pact was built to solve a coordination problem between teams who both answer to the same release calendar. Applying it to a third-party API like Stripe or GitHub breaks that premise at the root, because the provider never participates in verification, and the resulting gap between consumer expectation and provider behavior is the subject of this piece.
Pact's design and the assumption it rests on
The mechanic is simple to describe and hard to appreciate until you see what it prevents. The two services never run at the same time. There's no live call between them during the test itself. What passes between them is a document, and the document is what gets checked twice, once from each side.
The value of this arrangement is visible long before anything breaks. And a change that would break a dependent consumer gets caught in CI, before it reaches a staging environment where the cost of finding it is higher.
Provider states extend this further, and they're worth understanding closely because they come up again later as the piece's central technical problem. A precondition like "user 42 exists" or "order 7 is pending" has to be true before a given interaction can be verified, and Pact treats that precondition as a first-class part of the contract. The provider's verification harness is the thing responsible for making that precondition real, by running setup code that actually creates user 42 or puts order 7 into a pending state before the interaction test executes. None of this works unless the provider team is in the room, metaphorically speaking, running that harness on a schedule they control.
The Pact Broker's can-i-deploy check turns all of this into an enforceable gate, blocking a deploy when the required verified results aren't published to the broker. It looks at a specific consumer version and a specific provider version and tells you whether they're compatible, based on verified results already published to the broker. That check is only meaningful when both sides have actually published those results. It's a release gate built on the assumption that both parties did their half of the work.
The premise that breaks when the provider is a third party
When the provider is Stripe, or GitHub, or Slack, or any API outside the organization's control, the foundational assumption of the whole system collapses. A developer's test suite generates a contract file that Stripe will never see. Stripe will never run a verification suite against it. Stripe will never publish a result to a Pact Broker, because Stripe doesn't know the broker exists. The verification half of the loop gives Pact its authority, and it simply does not execute.
This is a boundary built into what the tool was designed to coordinate: two teams, both reachable, both willing to run verification on their own side of the contract. A third-party API provider is not a team a developer can onboard into that workflow. No amount of diligence on the consumer side produces a provider verification step, because that step requires the provider's cooperation by definition, and the provider was never asked and would have no reason to agree.
What's left after the collapse is a consumer-side test that generates a contract file, and a Pact mock server that replays that file locally on demand. That's a real artifact, and it does real work during development, but no step in the process has ever run against the live API to confirm the contract reflects what that API actually does. The contract file becomes, in effect, a description of what the developer expects the API to do, checked only against the developer's own expectations. It's a closed loop, and a closed loop cannot detect the one failure mode that matters most: the real API doing something the contract didn't anticipate.
The Microsoft ISE team's documented experience with Pact illustrates this precisely, and it's a useful case because the team was candid about the tradeoff. That's a reasonable decision for a proof of concept. It also means the provider verification half of the contract testing loop never ran, by design, and the team's own account says so directly.
can-i-deploy goes quiet in this same scenario. With no provider verification results ever landing in the broker, there's nothing for the gate to check against, so it has no signal to act on and nothing useful to report.
How the consumer-only mock diverges from the real API
Once provider verification is out of the picture, the Pact mock server becomes a static configuration: when it sees a given request, it returns a given response, fixed at the time the contract was written. It has no memory of prior calls, no mechanism to enforce that the response still matches reality, and no live connection of any kind to what the actual provider does today. It will keep returning that same response indefinitely, whether or not it's still true.
Consider a flow that occurs in nearly every integration that matters: create a resource, fetch it back, update it, then fetch it again, expecting the update to be reflected. A Pact mock configured against this flow still just returns whatever hardcoded response was recorded for each step. It can't tell a developer whether the resource created in step one has any bearing on what gets returned in step three, because the mock was never built with a notion of state carrying across requests. It answers each call in isolation, which is exactly the wrong model for a flow defined by its sequence.
Now picture a concrete case: a consumer integrates against Stripe's API, records a contract for a subscription object, and ships a Pact suite built on that contract. None of these changes touch the consumer's Pact mock at all, because the mock was never wired to the live provider in the first place. There is no mechanism by which a change on Stripe's side could propagate into a failing assertion on the consumer's side. The consumer's test suite keeps passing, green across the board, while the actual integration quietly stops working in whatever way the field rename or header requirement happens to break.
What provider states expose about the statefulness problem
Provider states are Pact's answer to the problem of preconditions, and examining how they're meant to work shows exactly where that answer stops functioning. A state like "user 42 exists" or "order 7 is in a pending status" is only ever made real by the provider running setup code inside its own verification harness before the interaction test executes. That setup step is a piece of real work a cooperating team does on their own infrastructure. Against a third-party API, no such team is available to do it.
So the developer ends up simulating the provider state entirely from the consumer side, which in practice means hand-writing a response body that looks like what "user 42 exists" ought to produce. The developer is authoring both sides of an argument and then agreeing with themselves.
This is where statefulness stops being an abstract concern. A mock configured to return a 200 with a subscription object attached tells you nothing about whether the real API requires a resource to have been created first, enforces an ownership check tied to the authenticated user, or returns something different entirely if a prior delete has already occurred on that same record. A stateless mock has no way to represent any of those dependencies, because by construction it answers every call the same way regardless of what came before it.
Sequential flows are the ordinary shape of most real integrations, not an edge case. The response at any given step in that chain depends on what happened at the steps before it, and a mock with no memory of those prior steps simply cannot model the dependency. Provider states give the test the appearance of handling this, the vocabulary of preconditions and setup, but without a provider actually wiring up that precondition, the mechanism is cosmetic. It makes the test read as though it accounts for state, but the code running underneath it is just as stateless as it would have been without the feature.
Where Pact adds value with a third-party provider
Pact is still the right tool to reach for in this context, and it does deliver something specific. A consumer-side Pact test, even one that never gets verified by a provider, produces a machine-readable, versioned record of what the consumer expects from the API it depends on. That's a meaningfully better artifact than an undocumented mock scattered across test files or a comment in a wiki nobody updates.
The Microsoft ISE team's experience backs this up from the other direction: they found contract tests valuable for a proof of concept integrating downstream systems built by teams at very different levels of engineering maturity, because the contracts gave everyone a structured, shared way to express interface expectations even without a fully built integration environment behind them. That's a real benefit, independent of whether provider verification ever runs.
Pact's matchers help here too. A type matcher, a regex, a date pattern, these relax the brittleness of asserting an exact value, so a test doesn't fail just because a timestamp changed between runs. That reduces noise even in a one-sided setup where no provider is checking the other end.
The same consumer-driven model extends to message-based testing for asynchronous interactions, like Kafka topics or other event streams, and Pact OSS added V4 support for this in Pact-JS as of the May 2025 update. That capability is genuinely useful for internal event-driven systems where both the producer and the consumer of a message participate in verification. It carries the identical constraint as the request-response case: the benefit depends on both sides showing up, and a third-party event source is no more likely to run a verification suite than a third-party REST API is.
The honest accounting is this: Pact run against a third-party API confirms that the consumer has encoded its own expectations correctly and consistently. It cannot confirm that those expectations match what the real API delivers now, or will still deliver after the next release the provider ships without telling anyone.
Filling the gap Pact leaves: stateful, continuously verified simulation
The two things missing from a consumer-only Pact mock are exactly the two things a replacement needs to supply. The first is memory across calls, so that a sequence like create, fetch, update, fetch again produces responses that actually reflect what happened in the earlier steps rather than fixed, disconnected answers. The second is continuous verification against the live provider, so that the simulator's behavior doesn't quietly drift away from reality the way a static mock inevitably does.
State matters because the logic that fails in production is almost never a single isolated call. It's a chain: create a payment intent, confirm it, fetch the resulting charge, issue a refund against it. Each step's correct response depends on the effects of the steps before it, and nothing stateless can represent that dependency chain with any fidelity.
Continuous verification matters for a separate reason: third-party APIs change on their own schedule, without consulting any particular consumer first. A simulator that gets checked against the live API before every release catches a field rename or a new required header while it's still sitting in a developer's local environment or a CI run, rather than after it has already reached a customer in production.
PactFlow's Drift tool, launched in March 2026, is a relevant development here and deserves a fair account. For a team already using PactFlow, that's a real addition to the toolkit. It's a CI-time check rather than something running continuously against production traffic, and it doesn't supply a stateful, locally hosted stand-in for the real service that developers can run requests against during ordinary day-to-day work. It solves a piece of drift detection. It does not solve the problem of statefulness.
Deciding what role Pact should play in an integration testing stack
The question to ask is not whether Pact works. What matters is who controls the provider. If the provider is an internal team willing to run verification and publish results to a broker, Pact fits the job precisely as designed, and the can-i-deploy gate carries real weight. If the provider is a third-party API, the verification loop never executes no matter how the consumer side is configured, and Pact by itself leaves a gap no amount of careful test-writing on the consumer's end can close.
For integrations with an external provider, the sensible stack layers tools according to what each one can actually promise. Consumer-side Pact tests still earn a place, documenting and versioning the team's expectations of the API in a form other engineers can read and trust. A narrow set of live calls against the real API, run in a limited pre-production stage, covers what no simulator can responsibly replicate: authentication edge cases, and how rate limiting actually behaves under production-level load.
Fault injection belongs in that middle layer as a core requirement, not an extra. A simulator that can return a 429, a 503, or a deliberate latency spike on command lets a team test its own retry and backoff logic against failure conditions that actually occur in the wild. A static Pact mock, built to return a 200, has no way to exercise that code path.
can-i-deploy still earns its keep in a mixed stack, wherever internal services participate in contracts verified through the broker. A simulator checked against the real API yesterday is a far stronger signal to deploy on than a mock last updated the day the integration was first written.

