JOEY VICTORINO Independent Technical Judgment

Field Notes · Incident Command · 7 min read

Crisis leadership is the conversion of incomplete technical facts into decisions.

A serious technical crisis runs two clocks at once. The evidence clock governs how quickly facts can be established: acquisition, analysis, and verification take the time they take. The decision clock is set by obligations, operating pressure, and stakeholders, and it does not wait for the first clock. The leadership work is managing the mismatch: partitioning what is known from what is inferred and unknown, sequencing decisions by reversibility rather than by volume of noise, and stating confidence in a form that can later go down without anyone losing credibility.

The two clocks

Most descriptions of incident management treat time as a single dimension: move faster. In practice two independent schedules are running, and confusing them produces most of the avoidable damage.

The evidence clock is governed by physics and process. Systems must be identified and acquired. Logs must be collected from places with different retention. Analysis produces findings in an order determined by what was available, not by what is most wanted. Some questions resolve in hours; some resolve in weeks; some never resolve, for reasons examined in Evidence Gaps Are Findings.

The decision clock is governed by other people's requirements. Contractual notification clauses start on defined triggers. Regulatory obligations impose their own deadlines: for SEC registrants, a material cybersecurity incident must be disclosed on Form 8-K generally within four business days of the materiality determination, and that determination itself must be made without unreasonable delay after discovery. Customers ask questions the moment they notice degradation. Operations force choices about whether systems stay up. None of these wait for analysis to finish.

The central skill is not accelerating the evidence clock, which is largely fixed once preparation has been set, but deciding well against the decision clock using whatever the evidence clock has produced so far, and doing it in a way that does not foreclose tomorrow's options.

Partition before you decide

The habit that makes this tractable is cheap and rarely practiced: every status update sorts its content into four bins, explicitly.

  • Established. Supported by direct evidence someone has examined. "This account authenticated from this address at this time, per these logs."
  • Inferred. A reasonable interpretation of established facts, with the inference visible. "The pattern is consistent with credential compromise rather than insider misuse, because X and Y." Inference is legitimate and must be labeled, because it is the category most likely to be revised.
  • Unknown, resolvable. Not yet answered, with a path and a rough time to an answer. This bin drives investigative priority.
  • Unknown, likely unresolvable. The telemetry does not exist, the retention expired, the system was rebuilt. Naming these early prevents the team from spending the crisis chasing an answer that is not available, and tells leadership which decisions will have to be made without that fact, permanently.

Partitioning does two things at once. It gives executives a defensible basis for choosing, because they can see which statements will hold. And it protects the responders, because an inference that was labeled as an inference can be revised without anyone appearing to have been wrong.

Sequence by reversibility, not by urgency

Under time pressure, the instinct is to rank decisions by how loudly they are being demanded. A better ordering principle is reversibility, because it maps directly onto how much evidence a decision deserves.

Reversible decisions can be made early on thin evidence and corrected. Isolating a segment, disabling an account, forcing a credential rotation, standing up additional monitoring: these cost something if wrong, and the cost is usually operational and recoverable. Under uncertainty, taking a reversible protective action and refining later is generally the stronger play.

Irreversible decisions deserve the evidence clock's output. Rebuilding a system before it has been preserved destroys the record. Telling customers that no data was affected cannot be unsaid; it can only be corrected, at a much higher price, and the correction is what converts an incident into a credibility event. Attribution stated publicly is close to irreversible. So is terminating an employee on a preliminary read of activity logs.

The practical rule follows: spend the crisis's scarce certainty on the irreversible decisions, and buy time for them where possible by taking reversible actions early. When an irreversible decision cannot wait, because a regulatory or contractual clock has run out, the correct move is to make it in language that matches the actual evidence rather than in language that sounds more reassuring. A notification that accurately describes what is known and what is still under investigation is defensible later. One that overstates confidence to sound calm is a liability that grows.

Confidence that is allowed to fall

Crisis communication inside the company should carry explicit confidence, and the organization should establish in advance that confidence can decrease. This sounds procedural and has a real effect: teams that cannot lower confidence without embarrassment will avoid raising it in the first place, which starves executives of usable assessments, or will defend an early assessment past the point the evidence supports it.

The formulation that works is plain: "We assess with moderate confidence that the exposure is limited to the staging environment. That assessment rests on the network flow data we have; if the missing authentication logs from the identity provider contradict it, it will change." Everyone in the room now knows the claim, its basis, and its failure condition. Compare that with "we believe it's contained," which is unfalsifiable, and which sets up the closure problem examined in the note on containment and closure.

Decide the rules before the crisis

A crisis is the worst available environment for inventing decision rules, because the people who should set them are tired, exposed, and receiving partial information. Most of the rules can be decided in advance, and preparation of this kind is precisely what incident response guidance means when it treats preparation as a phase rather than a document. Useful pre-commitments include who holds authority to declare an incident and at what threshold; who may take systems offline without further approval; what triggers engaging counsel and therefore what communications become privileged; who speaks to customers, regulators, and press, and who explicitly does not; the standing inventory of notification obligations, which is the artifact argued for in the note on enterprise security commitments; and what evidence is preserved before any rebuild.

Each of these, decided in calm conditions, removes a decision from the decision clock's queue and converts it into execution. That is the highest-leverage crisis work available, and it can only be done when there is no crisis.

The strongest objection

The credible objection is that this framing risks producing hesitancy: organizations in crisis need decisiveness, and a leader who narrates uncertainty can appear to be avoiding the call. The objection identifies a real failure mode, and the framing argues against it rather than for it. Partitioning is not deferral: its purpose is to identify precisely which decisions can be made now and to make them, while preventing the ones that should wait from being made by accident in a status meeting. A leader who says "we do not yet know how they got in; we are rotating credentials and isolating that segment now regardless; we will not make a statement about customer data until the flow logs are reviewed, which is Thursday" has made three decisions in one sentence. That is more decisive than confident vagueness, and it is auditable afterward.

A boundary condition worth stating: not every incident needs this apparatus. A contained commodity malware detection on a single workstation does not require partitioned status updates and reversibility analysis, and applying heavy process to routine events trains the organization to ignore it. The framing earns its cost when consequences are material and the facts are genuinely contested, which is also when organizations most often abandon structure in favor of urgency.

What good looks like

In a well-run crisis, several things are visible from outside the technical work. Status updates distinguish established facts from inferences without being asked. The same person is running the incident on day four as on day one, because roles were assigned in advance. Decisions are recorded with the evidence available at the time, so later review can distinguish a bad decision from a reasonable decision that met bad luck. Public and customer statements are narrower than the internal working theory, not broader. And when a fact arrives that contradicts an earlier assessment, the update says so directly, which is the behavior that makes every previous assessment worth believing.

Conclusion

The defining constraint of a technical crisis is that the evidence clock and the decision clock run independently, and leadership cannot make the first one match the second. What leadership can do is partition the facts honestly, spend scarce certainty on the decisions that cannot be taken back, take reversible protective action early, state confidence in a form that can be revised without collapse, and remove as many decisions as possible from the crisis by making them beforehand. The organizations that handle serious incidents well are rarely the ones that knew more sooner. They are the ones that knew what they did not know, and decided anyway, on the record.

When the technical question is contested and consequential

Special Situations is independent technical work for exactly these conditions: an incident where the explanation is disputed, a failure whose cause is contested, or a technical question that has to produce a defensible answer under time pressure. Scoped after intake, with expedited timelines available.

Know a leadership team that has never rehearsed who decides what during an incident? Send them this note while it is still theoretical.

Sources

← All Field Notes