pulses
LoadingNo completion record
DEVICE RECOVERY LAB / ADAM BATES
I built a recovery service for a simulated locker. I use controlled failures to test what it can establish before it resends a command.
PYTHON / HTTP / DURABLE JOURNALS
Observed by the test instrument; unavailable to recovery.
Loading
Loading captured evidence.
Select a snapshot to inspect it. Timestamps are observed; replay is slowed for inspection.
Inspect duplicate delivery, a lost reply, and reconnection.
02 / KNOW WHEN TO STOPSame record. Different reality.Two process crashes expose the limit of automatic recovery.
03 / TEST THE TRADE-OFFMore workers. Always better?Compare matched workloads, including the run that got slower.
A real run, replayed. This page does not execute the service or simulator. The drawing follows saved simulator observations; no physical hardware is involved. Run the processes locally ↗
I terminate the controller on either side of the simulated pulse. Both restarted journals contain unfinished intent. Only the separate test instrument reveals which action occurred.
LoadingNo completion record
LoadingNo completion record
A retry could repeat a completed action. Declaring success could invent one. I preserve uncertainty and require independent physical inspection.
Final captured outcomes. No real hardware is involved. The instrument is a separate store, excluded from recovery decisions; its pulse count is absent from the controller's command response.
I compare one recovery worker with four, using the same queued operations and faults injected by a separate HTTP proxy. The question is what happens to the unaffected work.
Loading measured workload evidence.
Lower typical delay. One important exception. Four workers finished faster in 8 of 9 mixed-fault pairs. The slowest four-worker batch took 8.009 s against 5.857 s for its one-worker pair. These observations do not isolate the cause of that variation.
I keep the contract, failure injection, recovery decisions and experimental results available alongside the interface.