DEVICE RECOVERY LAB / ADAM BATES

Did it happen?

I built a recovery service for a simulated locker. I use controlled failures to test what it can establish before it resends a command.

PYTHON / HTTP / DURABLE JOURNALS

The door opened. The reply didn’t.

RECORDED PROCESS OUTPUT
RECOVERY SERVICE DURABLE INTENT COMMAND → REPLY LOST LKR–01 / SIMULATED
DEVICE OBSERVATION—
—physical
action

Observed by the test instrument; unavailable to recovery.

SERVICE KNOWLEDGE Loading

Check before repeating.

Loading captured evidence.

  1. 01
    Preserve the requestLoading
  2. 02
    Ask the controllerLoading
  3. 03
    Decide from evidenceLoading
CAPTURED OUTCOMELoading evidence.
—command
sends
—journal
queries
FOLLOW THE EVIDENCELoading snapshots

Select a snapshot to inspect it. Timestamps are observed; replay is slowed for inspection.

A real run, replayed. This page does not execute the service or simulator. The drawing follows saved simulator observations; no physical hardware is involved. Run the processes locally ↗

GO DEEPER

Change the failure. Inspect the decision.

Download captured evidence ↓

Inspect the captured event timeline
    THE LIMIT OF AUTOMATIC RECOVERY02

    What if the record
    is missing too?

    I terminate the controller on either side of the simulated pulse. Both restarted journals contain unfinished intent. Only the separate test instrument reveals which action occurred.

    A / CRASH BEFORE THE PULSE
    —physical
    pulses
    CONTROLLER JOURNALLoadingNo completion record
    B / CRASH AFTER THE PULSE
    —physical
    pulse
    CONTROLLER JOURNALLoadingNo completion record

    The service has to stop.

    A retry could repeat a completed action. Declaring success could invent one. I preserve uncertainty and require independent physical inspection.

    Read the decision ↗

    Final captured outcomes. No real hardware is involved. The instrument is a separate store, excluded from recovery decisions; its pulse count is absent from the controller's command response.

    THE CONTROLLED EXPERIMENT03

    Does one slow device
    hold up the others?

    I compare one recovery worker with four, using the same queued operations and faults injected by a separate HTTP proxy. The question is what happens to the unaffected work.

    ACTUAL LOCAL MEASUREMENTS

    Completed operations over time

    1 worker4 workers

    Loading measured workload evidence.

    Lower typical delay. One important exception. Four workers finished faster in 8 of 9 mixed-fault pairs. The slowest four-worker batch took 8.009 s against 5.857 s for its one-worker pair. These observations do not isolate the cause of that variation.

    OPEN THE IMPLEMENTATION

    Every claim has
    a place to look.

    I keep the contract, failure injection, recovery decisions and experimental results available alongside the interface.

    Capture provenance & playback timing