A feature can pass every unit test and still fail the moment it meets the real world. The API sends a field in a slightly different format, the database behaves differently from an in-memory replacement, or a payment provider takes ten seconds to respond instead of one.
These are exactly the problems integration testing is designed to catch.
Integration testing across APIs, databases, and external services verifies whether separate parts of a software system actually work together. Instead of checking one function in isolation, the test exercises boundaries where data is serialized, stored, transmitted, transformed, or returned.
For modern applications, those boundaries are everywhere. A single checkout request might touch an HTTP API, PostgreSQL, a message broker, fraud detection, inventory, and a third-party payment processor.
The challenge is getting realistic confidence without building slow and fragile test environments. Strong integration strategies combine real infrastructure where it matters, controlled service doubles where it does not, contract testing, deterministic data, and carefully simulated failures.
Define Exactly What Your Integration Test Covers
The phrase “integration test” can mean very different things.
One team may use it for a repository test running against PostgreSQL. Another may mean a full environment containing fifteen microservices and several cloud systems.
Both can be valid, but unclear scope creates confusing test suites.
A useful approach is to distinguish narrow integration tests from broad system-level tests.
Narrow tests focus on one external boundary. For example, start a real database, call the application’s persistence layer, save a customer, and verify the actual stored record.
Martin Fowler’s testing guidance similarly describes integration tests as checks around boundaries such as databases, REST services, queues, filesystems, and other external collaborators.
This narrow scope gives faster feedback and makes failures easier to diagnose.
If a database integration test fails, engineers immediately know which boundary deserves attention instead of investigating an entire distributed system.
Test Databases With Real Database Technology
Mocking a database can be useful for unit tests, but it cannot prove that your SQL, schema, transaction behavior, constraints, or ORM mappings actually work.
A common mistake is replacing PostgreSQL with a lightweight in-memory database during testing.
The test becomes faster, but behavior may differ in areas such as SQL syntax, indexing, data types, JSON handling, locking, or transactions.
For important persistence code, test against the same database technology used in production.
Testcontainers is designed for exactly this scenario. It can start short-lived containers containing real databases, message brokers, and other infrastructure, giving each test environment a known state without requiring a permanently shared test server.
For example, a test can launch PostgreSQL, apply production migrations, create a customer, commit the transaction, and query the record again.
That verifies far more than whether a mocked repository returned the expected object.
Test Migrations Too
Database migrations deserve integration coverage.
A query can work against a freshly generated schema while failing on an upgraded production database.
Include migration execution in representative tests, particularly when changing columns, indexes, relationships, defaults, or constraints.
Database compatibility is part of the application contract.
Test API Boundaries, Not Just Business Logic
API integrations fail in surprisingly ordinary ways.
A header disappears. A date changes format. A nullable field becomes mandatory. Authentication tokens expire differently than expected.
Integration tests should therefore exercise real HTTP serialization and deserialization whenever practical.
If Service A calls Service B, test the client responsible for that communication rather than mocking the client itself.
You might verify that the application sends the expected authorization header, correctly serializes the request, handles a successful response, and interprets known error codes.
For internal services, contract testing can add another layer of protection.
Pact defines contract tests around the shared expectations between consumers and providers. Instead of deploying an entire ecosystem for every build, consumers document the interactions they depend on and providers verify that they continue satisfying those contracts.
That is particularly useful for independantly deployed microservices.
A provider can add new fields without breaking anyone, while removing a response field required by an existing consumer can be detected before deployment.
Use Service Stubs for External APIs
Third-party services create a different problem.
You usually cannot start a local copy of a bank, shipping carrier, identity provider, or payment network every time CI runs.
Calling the real production service is even worse.
Tests may create real transactions, pollute logs, hit rate limits, become dependent on network availability, or fail because the provider has an outage.
Instead, use controlled stubs for most integration scenarios.
WireMock, for example, can return predefined HTTP responses based on request matching and can also capture requests for later verification.
Suppose your application communicates with a delivery API.
You can configure the stub to return a normal shipment response, an expired token, a rate-limit error, or malformed data.
This gives repeatable behavior without depending on the external provider’s availability.
For services that offer dedicated sandbox environments, keep a smaller number of tests against that sandbox as an additional compatibility check.
The stub protects daily development speed. The sandbox verifies that your assumptions still match reality.
Test Failures That Are Hard to Reproduce Naturally
Happy-path integration testing is not enough.
Production systems fail in inconvenient ways.
A service may respond slowly rather than returning an immediate error. Connections can reset midway through responses. A server might return HTTP 500 for thirty seconds and then recover.
These situations are often difficult to trigger reliably against a real external service.
Service virtualization makes them deterministic.
WireMock supports injected delays, random latency distributions, slow chunked responses, empty responses, malformed data, and connection-reset behavior.
This makes it possible to test questions such as:
Does the application time out correctly? Does it retry? Is exponential backoff working? Does a timeout accidentally create the same order twice?
These are often more important than the successful request itself.
A system that works when every dependency is healthy is easy to build. A system that behaves predictably when dependancies fail requires deliberate testing.
Treat Webhooks as Two-Way Integrations
External integrations are not always initiated by your application.
Payment systems, Git platforms, messaging providers, and other services often communicate back through webhooks.
Webhook testing deserves its own strategy because incoming events can be delayed, duplicated, delivered out of order, or retried.
GitHub’s official webhook guidance, for example, supports testing whether events produce webhook deliveries and forwarding webhook traffic to a local development environment so handlers can be tested before production deployment.
Your tests should verify more than simply returning HTTP 200.
Check signature validation, event parsing, duplicate delivery handling, unknown event types, and retry safety.
Imagine a payment-completed webhook arriving twice.
If the handler processes both independently, a customer might receive duplicate credits or inventory could be decremented twice.
Webhook consumers should generally be designed to be idempotent, and integration tests are an excellent place to verify that behavior.
Keep Test Data Deterministic and Isolated
Shared test environments often become unreliable because everyone manipulates the same data.
One CI job creates Customer A. Another deletes it. A third assumes the account already exists.
Suddenly, tests fail depending on execution order.
Good integration tests should control their own state.
Create unique records for each run, reset containers between suitable test groups, or use transactions and cleanup mechanisms.
Testcontainers helps by creating disposable infrastructure that can start from a predictable state rather than relying on a database that has accumulated months of test records.
Isolation also makes parallel execution much safer.
A test should not depend on another test having run previously.
When cleanup fails, design test data so collisions still remain unlikely. Unique identifiers, generated namespaces, or run-specific schemas can help.
Deterministic test data improves both execution speed and reliablity.
Avoid Turning Integration Tests Into End-to-End Tests
Integration tests become slow when their scope quietly expands.
A repository test starts the database. Then someone adds Redis. Later it also launches Kafka, authentication, inventory, and four downstream services.
Eventually, every test requires half the production architecture.
That defeats one of the main advantages of focused integration testing.
Use real dependencies only when they are part of the boundary currently being verified. Replace unrelated components with stable doubles.
Testing guidance for microservice architectures similarly recommends focused gateway and persistence integration tests, with broader component or end-to-end tests used at higher layers.
For example, while testing your PostgreSQL repository, the payment provider does not need to exist.
While testing your Stripe-style gateway adapter, your production database probably does not need to exist either.
One boundary at a time creates faster tests and clearer failures.
Run Integration Tests Strategically in CI
Integration tests are usually slower than unit tests because they start infrastructure or communicate over networks.
That does not mean they should run only once before release.
Organize CI around feedback speed.
Fast unit tests and static checks can run first. Focused integration tests follow immediately afterward. Larger sandbox or full-system checks may run later, nightly, or during release pipelines.
This avoids making developers wait for expensive environments when basic code already fails.
Containerized infrastructure also makes CI environments more reproducable because databases and supporting systems can be provisioned on demand rather than manually maintained.
Testcontainers supports databases, message brokers, browsers, and many other containerized dependencies across multiple programming languages.
Track integration-test duration and flaky-test rates over time.
If the suite gradually becomes slower or less predictable, treat that as an engineering problem instead of accepting permanent test debt.
Integration testing provides confidence in the places where real software often breaks: boundaries.
Test actual database behavior where persistence matters, exercise HTTP serialization around APIs, protect internal service relationships with contract tests, and use controlled stubs for external providers.
Failure simulation should cover latency, malformed responses, network errors, retries, and duplicate events – not only successful requests.
Keep every test isolated and make its scope obvious. A focused integration test should explain exactly which relationship it protects.
Start by mapping the external boundaries in one important application workflow. Identify the API calls, database operations, webhooks, queues, and third-party services involved, then make sure each critical boundary has a reliable test.
That usually produces more confidence than simply adding another large end-to-end test.

