Simulate before you integrate
Some systems are easy to test: call a function, check the result. Others only come alive when the real world feeds them: machines producing events, upstream systems pushing orders, sensors streaming readings at thousands of messages a minute. You can't book a production line for a Tuesday afternoon just to run your tests.
That's why so much of my tooling work is simulation: software that behaves like the real environment closely enough that the platform consuming it can't tell the difference. Done well, a simulator becomes the most valuable test asset a team owns.
What a good simulator gives you#
- End-to-end tests without the hardware. The whole pipeline runs against realistic input, on any laptop or CI runner.
- Performance testing on demand. Turn the speed up to ten times a real shift and watch where things queue.
- Failures on purpose. Real systems send duplicates, drop messages and deliver them late. A simulator can do all of that deliberately and repeatably.
- Replay. Record a scenario once, replay it after every change and compare results.
- Demos that don't lie. Stakeholders see the product working on data that looks real, because it behaves like real data.
Rule one: reproducibility#
A simulator that produces different data every run turns every failing test into a mystery. Seed the randomness and the same scenario always produces the same stream:
import random
from dataclasses import dataclass
from datetime import datetime, timedelta, timezone
@dataclass(frozen=True)
class Event:
seq: int
source: str
timestamp: datetime
value: float
def generate(sources, start, interval, seed=42):
rng = random.Random(seed) # a private RNG: nothing else can disturb the sequence
seq, now = 0, start
while True:
for source in sources:
seq += 1
yield Event(seq, source, now, round(rng.gauss(50.0, 4.0), 2))
now += interval
It's a generator, so it can run forever without holding anything in memory and itertools.islice takes exactly as
many events as a test needs.
Rule two: simulated time, not wall-clock time#
Tie the simulator to the real clock and a one-hour scenario takes an hour. Generate timestamps from a simulated clock instead, then decide separately how fast to emit them: real time for a demo, as fast as possible for a load test:
import time
def paced(events, speed=1.0):
"""Emit events at `speed` times real time. speed=None means as fast as possible."""
first_sim = first_wall = None
for event in events:
if speed:
first_sim = first_sim or event.timestamp
first_wall = first_wall or time.monotonic()
due = first_wall + (event.timestamp - first_sim).total_seconds() / speed
time.sleep(max(0.0, due - time.monotonic()))
yield event
Note time.monotonic(): the wall clock can jump (NTP corrections, daylight saving); a monotonic clock can't.
Rule three: inject the failures production will send#
The happy path is the least interesting thing to test. Wrap the clean stream in a layer that misbehaves in controlled, seeded ways:
def with_faults(events, seed=7, duplicate=0.01, drop=0.005, delay=0.02):
rng = random.Random(seed)
held = []
for event in events:
roll = rng.random()
if roll < drop:
continue # lost in transit
if roll < drop + delay:
held.append(event) # will arrive late, out of order
continue
yield event
if rng.random() < duplicate:
yield event # delivered twice
if held and rng.random() < 0.5:
yield held.pop(0)
yield from held
Now the consumer's deduplication, ordering and gap detection get exercised on every run, long before a real network does it at 2 a.m.
Putting it together#
from itertools import islice
import json
start = datetime(2026, 1, 1, 6, 0, tzinfo=timezone.utc)
stream = with_faults(generate(["line-1", "line-2", "line-3"], start, timedelta(seconds=1)))
with open("scenario.jsonl", "w") as out:
for event in islice(paced(stream, speed=None), 10_000):
out.write(json.dumps({**event.__dict__, "timestamp": event.timestamp.isoformat()}) + "\n")
Swap the file for a message queue, an HTTP endpoint or a database and the same generator drives a real integration test.
Beyond the basics: what makes simulators hard#
A generator of plausible-looking events is where most simulators stop. The ones that genuinely de-risk a platform go further in five directions.
1. Model the domain, not just the data#
Random values that happen to have the right types will pass a schema check and still be nonsense: an order completed before it was released or a batch cancelled twice. Real systems produce events that follow a lifecycle, so the simulator should model that lifecycle explicitly as a state machine.
import random
TRANSITIONS = {
"created": ["released", "cancelled"],
"released": ["in_production", "cancelled"],
"in_production": ["completed", "on_hold"],
"on_hold": ["in_production", "cancelled"],
}
def lifecycle(order_id, rng):
state = "created"
yield order_id, state
while state in TRANSITIONS:
state = rng.choice(TRANSITIONS[state])
yield order_id, state
rng = random.Random(42)
for order_id, state in lifecycle("PO-1001", rng):
print(order_id, state)
Every generated sequence is now valid by construction. When you do want invalid sequences (to test how the consumer rejects them), inject them deliberately as a named fault, just like duplicates and delays.
2. Respect backpressure#
A simulator that pushes events faster than the system under test can accept doesn't test the system. It floods it and hides the real bottleneck. Put a bounded queue between the generator and the sink, so a slow consumer slows the producer down and you can measure how far behind it falls:
import asyncio
async def run(events, sink, rate, max_in_flight=100):
"""Emit events at `rate` per second into a bounded queue and report the peak backlog."""
queue = asyncio.Queue(maxsize=max_in_flight)
peak_backlog = 0
async def produce():
nonlocal peak_backlog
for event in events:
await queue.put(event) # blocks whenever the consumer falls too far behind
peak_backlog = max(peak_backlog, queue.qsize())
await asyncio.sleep(1 / rate) # the simulated arrival rate
await queue.put(None)
async def consume():
while (event := await queue.get()) is not None:
await sink(event)
await asyncio.gather(produce(), consume())
return peak_backlog
A consumer that keeps up holds the backlog near zero. One that can't will climb until the backlog reaches
max_in_flight, at which point the producer is throttled to the consumer's real speed. That peak is the clearest signal
you'll get that the consumer cannot sustain the rate you asked for. Pacing the producer matters: without the sleep,
the producer would fill the queue before the consumer ever ran and even an instant consumer would look overloaded.
3. Record, replay and compare with golden files#
Seeding makes a scenario reproducible; recording it makes it portable. Save every emitted event, then replay the same file against each new build and compare the consumer's output with a known-good "golden" result. Any difference is either a bug or an intended change that needs a new golden file.
import json
from pathlib import Path
def record(events, path):
with Path(path).open("w") as out:
for event in events:
out.write(json.dumps(event, sort_keys=True) + "\n")
def replay(path):
with Path(path).open() as source:
for line in source:
yield json.loads(line)
def matches_golden(actual, golden_path):
expected = json.loads(Path(golden_path).read_text())
return actual == expected
A library of recorded scenarios (a normal shift, a network outage, the worst peak of last year) becomes a regression suite that runs in CI and protects the platform for years.
4. Keep the event schema a shared contract#
The simulator is only useful while it speaks exactly the same language as the real systems. If the real schema gains a field and the simulator doesn't, every test is quietly testing the past. Define the schema once, version it and have both the consumer and the simulator validate against it.
from dataclasses import dataclass, fields
SCHEMA_VERSION = 2
@dataclass(frozen=True)
class OrderEvent:
schema: int
order_id: str
state: str
timestamp: str
def validate(payload):
expected = {f.name for f in fields(OrderEvent)}
if set(payload) != expected:
raise ValueError(f"fields {sorted(set(payload) ^ expected)} do not match the contract")
if payload["schema"] != SCHEMA_VERSION:
raise ValueError(f"schema {payload['schema']} is not the supported version {SCHEMA_VERSION}")
return OrderEvent(**payload)
A contract test that validates a sample of simulated events on every build catches drift the day it happens, not the day a customer's system sends the new field.
5. Load profiles and latency percentiles#
Averages hide exactly the behaviour you care about. Drive the platform through profiles that mirror reality (a normal shift, a sustained peak, a sudden spike after an outage) and report latency as percentiles:
import statistics
PROFILES = {
"normal_shift": [(8 * 3600, 20)], # (duration in seconds, events per second)
"peak": [(3600, 20), (1800, 120), (3600, 20)],
"recovery_spike": [(600, 0), (120, 500), (1800, 20)],
}
def latency_report(latencies_ms):
cuts = statistics.quantiles(latencies_ms, n=100)
return {"p50": cuts[49], "p95": cuts[94], "p99": cuts[98], "max": max(latencies_ms)}
The p99 tells you what your unluckiest users experience. The recovery spike tells you whether the platform can catch up after an outage or will fall further behind forever. Those two numbers have saved more go-lives than any average ever has.
The payoff#
A simulator built this way becomes one of the most valuable assets a team owns. It enables end-to-end tests without hardware, performance tests on demand, deliberate chaos, regression suites of recorded scenarios and demos that behave like the real thing. Every one of them grows from the same foundations: seeded, simulated-time, deliberately faulty, domain-accurate and contract-bound. They pay for themselves the first time a bug is caught on a laptop instead of on a line.
