After a Production Incident

Trace how dependencies, detection, recovery, and change interacted. Preserve the timeline and uncertainty before choosing an explanation.

Questions worth answering

  • What is directly supported by the timeline?
  • Which failure modes remain plausible?
  • What would distinguish competing explanations?

The outcome to protect

Reduce repeat exposure. An untested explanation or a closed ticket does not prove resolution.

How we investigate

Reconstruct evidence, distinguish facts from hypotheses, review detection and recovery gaps, and design the next discriminating test.

Decisions & outcomes

  • Evidence-linked incident timeline
  • Failure modes and adjacent exposure
  • Actions with closure and reverification criteria

Trace the missing notification without guessing the cause

A fictional scenario, not a customer result.

The question to resolveWas a completed renewal followed by one durable notification, or did the handoff lose work?

  1. Renewal committed
  2. Event handed off
  3. Notification worker
  4. Delivery receipt

Observed record

The reconstructed timeline contains a committed renewal and a missing notification receipt. It does not locate where the handoff failed.

Evidence basis
Observed
Verification result
Partially verified
Freshness
Fresh

Hypothesis

An interruption between committing the renewal and publishing its event could explain the symptom. That explanation remains an inference until discriminating evidence is gathered.

Evidence basis
Inferred
Verification result
Unverified
Freshness
Unknown

The resulting decision

Preserve the relevant timeline, test the handoff hypothesis in a safe environment, and give the resulting corrective action a measurable closure condition.

Scope of this example

This is a fictional incident review. A plausible explanation is not a confirmed cause, and an action list alone does not verify prevention.

Choose one critical flow.

Tell us what must work and what you need to know. We agree the scope before delivery.