<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Ultimate TTX</title><link href="https://ultimatettx.com/blog/" rel="alternate"/><link href="https://ultimatettx.com/blog/feeds/all.atom.xml" rel="self"/><id>https://ultimatettx.com/blog/</id><updated>2026-09-15T00:00:00-06:00</updated><subtitle>Notes on running tabletop exercises that find things.</subtitle><entry><title>Which roles to put in the room</title><link href="https://ultimatettx.com/blog/2026/which-roles-to-put-in-the-room/" rel="alternate"/><published>2026-09-15T00:00:00-06:00</published><updated>2026-09-09T00:00:00-06:00</updated><author><name>Kenneth Ingham</name></author><id>tag:ultimatettx.com,2026-09-15:/blog/2026/which-roles-to-put-in-the-room/</id><summary type="html">&lt;p&gt;Every role added to an exercise improves coverage of the incident and costs half a day or more of somebody's working time. The scenario should decide the roster, and the library shows how much it varies.&lt;/p&gt;</summary><content type="html">&lt;p&gt;The planning conversation reaches the participant list, and it is usually
treated as an invitation problem. It is not. It is the decision that determines
what the exercise is capable of finding, and it has a real cost on the other
side.&lt;/p&gt;
&lt;h2&gt;What another role buys&lt;/h2&gt;
&lt;p&gt;An incident does not stay inside one team. It starts somewhere, is recognized
somewhere else, and its consequences land in places that have nothing to do
with computers.&lt;/p&gt;
&lt;p&gt;A role that is missing from the room does not merely go unrepresented. Its part
of the response gets &lt;strong&gt;assumed&lt;/strong&gt;, and assumptions in an exercise are generous
in a way that real people are not.&lt;/p&gt;
&lt;p&gt;The first thing to get right is that "IT" is not a role. In most organizations
of any size it is several groups with different tools, different on-call
rotations and different managers, and the incident crosses all of them:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Network operations.&lt;/strong&gt; Frequently the group that can actually stop lateral
  movement while it is happening, by taking segments apart. If containment is
  going to be tested at all, somebody who can do it has to be present.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Server administration&lt;/strong&gt;, which owns the systems being restored and the real
  numbers about how long that takes.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Endpoint administration&lt;/strong&gt;, which owns the machines where most incidents
  begin.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Security&lt;/strong&gt;, and in larger organizations security operations as a separate
  function with its own tooling and its own queue.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The helpdesk&lt;/strong&gt;, which is where most incidents are first visible, usually as
  something that does not look like an incident.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Identity administration&lt;/strong&gt;, which holds the accounts and the authority to
  disable them.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Compliance&lt;/strong&gt;, where it exists as its own group rather than as somebody's
  additional duty.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Then the parts of the organization that are not IT at all, and where the
exercise usually finds its most durable material:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Management&lt;/strong&gt;, which holds the decisions that cost money. A room without
  somebody who can commit the organization stalls at the first real choice.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Contracts.&lt;/strong&gt; Worth separating two questions that sound like one. &lt;em&gt;What data
  is on this system&lt;/em&gt; is usually answerable in the room: the project manager
  knows, and so does whoever uses the system daily, which is one of the better
  arguments for having an actual user present rather than only their manager.
  &lt;em&gt;What reporting obligations attach to that data&lt;/em&gt; is a different question, and
  contracts is where it lives. That one stalls a room, and it is worth building
  into a scenario deliberately.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Legal&lt;/strong&gt;, for the questions the contracts do not settle.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Communications.&lt;/strong&gt; Somebody has to write what is said to customers, staff
  and regulators, and doing that badly is its own incident. &lt;a href="https://www.cisa.gov/sites/default/files/2026-09/joint-guidance-communicating-under-pressure-508c.pdf"&gt;Joint guidance
  from CISA, the FBI and four partner
  agencies&lt;/a&gt;
  recommends designating an incident lead, a communications lead and a
  spokesperson with predefined approval paths, and recommends exercising that
  structure through tabletop exercises specifically. Where a plan names those
  roles and no exercise ever tests them, the plan is untested at its most
  visible point.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Human resources&lt;/strong&gt;, for anything involving a person rather than a system.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Finance&lt;/strong&gt;, for emergency spending and for the question of what the outage
  is costing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Physical security&lt;/strong&gt;, wherever premises, media or access badges are
  involved.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The application owner&lt;/strong&gt;, who is the only person who knows what a system is
  actually used for and therefore whether an outage is survivable.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;If the organization uses a managed service provider, they should be in the
room.&lt;/strong&gt; For many small organizations the provider &lt;em&gt;is&lt;/em&gt; the response capability,
and an exercise that models their part rather than involving them is exercising
a guess. Their contract is also the thing that determines how fast they arrive,
which is worth discovering in a planning conversation rather than at two in the
morning.&lt;/p&gt;
&lt;h2&gt;What a missing role actually costs, and one it does not&lt;/h2&gt;
&lt;p&gt;The value of a role is not uniform, and it is worth being precise rather than
arguing that everyone should attend.&lt;/p&gt;
&lt;p&gt;Take reporting obligations, which are often used to argue for legal counsel.
For a defense contractor this is one of the clearer areas rather than one of
the murkier ones. The clause is
&lt;a href="https://www.law.cornell.edu/cfr/text/48/252.204-7012"&gt;DFARS 252.204-7012&lt;/a&gt;, and
it is unusually specific about all three. It defines what counts as a reportable
cyber incident. It requires the contractor to &lt;em&gt;rapidly report&lt;/em&gt; one to DoD, which
paragraph (a) defines as &lt;strong&gt;within 72 hours of discovery&lt;/strong&gt;. And it requires,
separately, that images of the affected systems and the relevant monitoring data
be preserved for &lt;strong&gt;at least 90 days&lt;/strong&gt; from the report, so that the department can
ask for the media or decline it. The reporting portal is dibnet.dod.mil and can
be visited in advance, so an organization can know before any incident exactly
what it will be asked for.&lt;/p&gt;
&lt;p&gt;That makes the obligation a &lt;strong&gt;preparation&lt;/strong&gt; question rather than a legal one,
and the exercise finding is usually not "we needed a lawyer" but "nobody had
read it, and nobody had looked at the form." Both are fixable in an afternoon
by somebody who is already on staff.&lt;/p&gt;
&lt;p&gt;A related observation belongs in the report when it applies: an organization
that has been in the defense industrial base for years and has never filed a
report is not describing an unusually quiet decade. Laptops get stolen. People
click things. A clean history of that length more often means the reporting
threshold was never understood than that nothing ever crossed it.&lt;/p&gt;
&lt;p&gt;Contracts is the opposite case. "Which agreements cover the data on this
system" is not knowable from the IR plan, not knowable by IT, and not knowable
at half past five in the evening. That absence changes what the room can do.&lt;/p&gt;
&lt;h2&gt;What another role costs&lt;/h2&gt;
&lt;p&gt;Half a working day, at least, once the pre-exercise briefing is counted. That
is the honest number, and it is per person.&lt;/p&gt;
&lt;p&gt;For a small organization this is not a rounding error. Six people out for half
a day is a meaningful piece of a week's capacity, and the people whose presence
would help most are frequently the hardest to release, because they are the
same people holding operations together while the exercise runs.&lt;/p&gt;
&lt;p&gt;What that cost is &lt;em&gt;not&lt;/em&gt;, in practice, is travel. Organizations do not fly
people in for a tabletop exercise, and they are right not to: with travel
either side it becomes a three-day commitment plus expenses, for half a day of
participation. Where somebody is not local, having them attend remotely is
nearly as good and costs a fraction as much. A remote participant is a slightly
diminished participant rather than an absent one, and that trade is almost
always better than dropping the role.&lt;/p&gt;
&lt;p&gt;There is a second cost that is easier to miss. A room that is too large stops
being a discussion. Participants with no part in the current inject disengage,
and a disengaged participant is not merely idle. They are reading email, and
the injects continue while they do. When the discussion reaches something they
could have contributed to—and it does, because that is why they were
invited—they have missed the moment. The exercise then records a gap that was
not really a gap, or misses one that was.&lt;/p&gt;
&lt;p&gt;A room that is too large also changes character: the people still engaged start
performing for an audience rather than working the problem.&lt;/p&gt;
&lt;p&gt;So more roles is not monotonically better, even ignoring the money.&lt;/p&gt;
&lt;h2&gt;Deciding rather than inviting&lt;/h2&gt;
&lt;p&gt;The useful reframing is that the scenario determines the roster, not the org
chart. This is not a principle to be taken on trust—it is visible in any
scenario library that has been built deliberately, and it is worth looking at
before planning a particular exercise.&lt;/p&gt;
&lt;p&gt;The scenario library Ultimate TTX builds for subscribers shows the pattern
plainly. Across those scenarios three roles appear almost everywhere:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;ransomware with exfiltration&lt;/strong&gt; scenario reaches ten roles, including the
  helpdesk, backup administration and endpoint administration, because the
  response spans first contact, containment, restoration and disclosure.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;business email compromise&lt;/strong&gt; scenario reaches five, and finance is central
  to all of them. Backup administration has no part in it at all.&lt;/li&gt;
&lt;li&gt;An &lt;strong&gt;insider data theft&lt;/strong&gt; scenario brings in human resources and drops IT
  operations.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;physical access and removable media&lt;/strong&gt; scenario needs physical security,
  which appears nowhere in the financial or identity scenarios.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;service provider compromise&lt;/strong&gt; scenario needs network operations and the
  application owner, and much of it is about somebody else's environment.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Build one standing invitation list and reuse it, and some exercises carry
passengers while others are missing the person who mattered. Read the scenario
first; it will tell you who has to be there.&lt;/p&gt;
&lt;p&gt;Two practical devices help.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Name the role, not the person.&lt;/strong&gt; The question is which functions have to be
represented; who fills each one is a scheduling problem after that. It also
makes substitution visible, and substitution matters: a deputy who does not
hold the authority will answer as though they do.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Absence is a finding.&lt;/strong&gt; When a role cannot be released, the exercise still
runs, and the report records that the response was modeled without it. That is
more useful than quietly proceeding, because "we could not free the only person
who can authorize an emergency shutdown" deserves management's attention, and
it recurs during real incidents for exactly the same reason.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;One role is not on the scenario's list, and it belongs in every room: whoever
would coordinate the response.&lt;/strong&gt; The scenario decides which functions the
incident touches; coordination is not one of those functions, it is the thing
that sequences the rest of them. Sending a deputy is worse here than anywhere
else on the roster, because the exercise cannot then observe whether the
arrangement works—it observes a stand-in improvising one. Why that role matters,
and what it has to not be doing, is a subject of its own.&lt;/p&gt;
&lt;p&gt;Communications is the other role worth arguing for beyond first principles.
Where a scenario requires the organization to say anything to anyone outside the
response team, the function that says it should be present rather than
represented by whoever is nearest—and the joint guidance &lt;a href="https://www.cisa.gov/sites/default/files/2026-09/joint-guidance-communicating-under-pressure-508c.pdf"&gt;Communicating Under
Pressure&lt;/a&gt;,
published by six agencies in September 2026, recommends exercising communication
plans through "simulated tabletop exercises" and naming communications teams
specifically.&lt;/p&gt;
&lt;h2&gt;The roster shapes the report, not just the day&lt;/h2&gt;
&lt;p&gt;Choosing the room is also choosing what the report can say, and it is worth
thinking about at planning time rather than discovering afterward.&lt;/p&gt;
&lt;p&gt;A finding about a role that was not represented is weaker, and it should be:
nobody from that function was there to explain what actually happens. The
report has to say so, and a reader is entitled to discount it. Conversely, a
finding that rests on two functions contradicting each other in the room—the
backup administrator and the person who would be woken up, say, or IT and
contracts—is among the strongest material an exercise produces, and it is only
available if both were present.&lt;/p&gt;
&lt;p&gt;The recommendations divide along the same lines. A finding whose fix is
political, requiring somebody senior to settle a disagreement between two
groups, lands very differently depending on whether both groups were in the
room and heard it emerge.&lt;/p&gt;
&lt;p&gt;This is also the strongest argument for management attending rather than being
briefed afterward. The recommendations needing an executive decision are
exactly the ones that arrive in a report as a surprise, and an executive who
watched the disagreement happen does not have to be persuaded that it is real.&lt;/p&gt;
&lt;h2&gt;The failure mode worth avoiding&lt;/h2&gt;
&lt;p&gt;The version that wastes everyone's time is the room built from whoever was
available, running a scenario chosen independently of them. It produces a
pleasant discussion, a report with thin findings, and a general impression that
tabletop exercises are a compliance ritual.&lt;/p&gt;
&lt;p&gt;Who is in the room is part of the scenario design. Choosing deliberately is
most of what separates an exercise that finds something from one that fills a
requirement.&lt;/p&gt;
&lt;p&gt;Ultimate TTX plans &lt;a href="/facilitated-booking/"&gt;facilitated exercises&lt;/a&gt; from the
scenario outward, including which roles have to be in the room for it to be
worth running.&lt;/p&gt;</content><category term="Facilitation"/><category term="roles"/><category term="planning"/><category term="participants"/><category term="scoping"/></entry><entry><title>"How long would this take you?"</title><link href="https://ultimatettx.com/blog/2026/how-long-would-this-take-you/" rel="alternate"/><published>2026-09-08T00:00:00-06:00</published><updated>2026-09-09T00:00:00-06:00</updated><author><name>Kenneth Ingham</name></author><id>tag:ultimatettx.com,2026-09-08:/blog/2026/how-long-would-this-take-you/</id><summary type="html">&lt;p&gt;Everyone estimates assuming things go well, and an incident is the situation where things are not going well. The exercise cannot produce the true number, but it can show that nobody has one.&lt;/p&gt;</summary><content type="html">&lt;p&gt;It is a regular question in an exercise, and it is worth asking often.
Restoring that server: how long? Finding which accounts the attacker used: how
long? Getting the replacement hardware: how long?&lt;/p&gt;
&lt;p&gt;The answers are optimiztic. Not occasionally, and not only from the
overconfident. They are optimiztic in the same direction, from careful people,
about work they have done before, and sometimes they are wildly optimiztic.&lt;/p&gt;
&lt;p&gt;The reason is not carelessness. &lt;strong&gt;Everybody estimates on the assumption that
things go well, and an incident is by definition the situation in which things
are not going well.&lt;/strong&gt; Skill offers no protection here and may work against it:
an experienced administrator's estimate is drawn from a real memory of doing
the work competently. During the incident those conditions are gone, and what
replaces them is worse than merely different: the same work, under time
pressure, in front of an audience, with consequences attached.&lt;/p&gt;
&lt;h2&gt;The estimate is the task; the incident is the conditions&lt;/h2&gt;
&lt;p&gt;The number is usually the work itself, measured on a good day, by someone who
has done it. What it leaves out is everything the incident puts around the
work.&lt;/p&gt;
&lt;p&gt;Start with the room. During an incident the office fills with people who need
answers: managers asking what is happening, end users asking when their systems
will be back, somebody stating—or shouting—what the downtime is costing per
hour. A colleague, part-way through an incident, had to tell his own boss and
several others to leave the area so that he and his team could concentrate.
That is a reasonable thing to have to do and a difficult thing to do, and the
time it takes to reach the point of doing it is not in anybody's estimate.&lt;/p&gt;
&lt;p&gt;It is also, as it turns out, the recommended practice. Joint guidance from
CISA, the FBI and the Australian, Canadian, New Zealand and UK cyber centers
tells organizations to run communications through designated liaisons,
scheduled syncs and a single intake path, explicitly in order to &lt;a href="https://www.cisa.gov/sites/default/files/2026-09/joint-guidance-communicating-under-pressure-508c.pdf"&gt;shield
engineers from external
interruptions&lt;/a&gt;.
Nobody should have to improvise that at the moment it is needed, and an
organization that has not arranged it in advance is relying on somebody being
willing to throw their own manager out of the room.&lt;/p&gt;
&lt;p&gt;The interruptions themselves cost more than people expect, and this has been
measured. In &lt;a href="https://doi.org/10.1145/1240624.1240730"&gt;a field study of ordinary computer
work&lt;/a&gt;, Iqbal and Horvitz logged what
happened after email and instant messaging alerts. Where somebody answered an
email alert immediately, an average of &lt;strong&gt;sixteen and a half minutes&lt;/strong&gt; passed
before they were working again in the same program and document they had left,
against an overall rate of close to four alerts an hour. In 27% of cases they
had still not returned to it two hours later.&lt;/p&gt;
&lt;p&gt;The authors are careful about what that measures, and it is worth repeating:
the clock stops when the person is back in the software they were using, which
is not the same as being back in the task. Getting the window in front of you
again is the first step of resuming work, not the completion of it. The real
cost of the interruption is therefore larger than the figure, not smaller.&lt;/p&gt;
&lt;p&gt;There is a mechanism behind that number, and it makes the incident case worse
rather than better. Leroy's work on &lt;a href="https://doi.org/10.1016/j.obhdp.2009.04.002"&gt;&lt;strong&gt;attention
residue&lt;/strong&gt;&lt;/a&gt; describes what
persists: cognitions about the first task that continue after somebody has
stopped working on it, switched to a second, and is now working on the second.
Those lingering thoughts are a cognitive load, and the resources they consume
are not available to the task actually in front of the person.&lt;/p&gt;
&lt;p&gt;What decides how much residue there is turns out to be whether the first task
was &lt;strong&gt;finished&lt;/strong&gt;. Across two experiments, people who were stopped mid-task
carried the unfinished one with them and did measurably worse on what came
next; people who finished first did better.&lt;/p&gt;
&lt;p&gt;That is the incident case exactly. Interruptions during an incident almost
never arrive at a clean stopping point—the interruption is what makes the task
unfinished, and the person carries the restore, or the log search, or the
half-formed theory about what happened, into the conversation with the director
and back out again.&lt;/p&gt;
&lt;p&gt;Leroy also found a way out that an incident cannot offer. Finishing the first
task under time pressure produced &lt;em&gt;less&lt;/em&gt; residue than finishing it at leisure,
because working against the clock narrows what somebody considers, which leaves
them more confident it is done and better able to stop thinking about it. Time
pressure helps only in combination with completion. During an incident there is
plenty of the first and very little of the second.&lt;/p&gt;
&lt;p&gt;So the measured figures should be read as a floor rather than a model. Ordinary
office work, interrupted at ordinary moments, is the benign case.&lt;/p&gt;
&lt;h2&gt;One task, several bottlenecks&lt;/h2&gt;
&lt;p&gt;Interruption stretches whatever the work turns out to be. The next problem is
that the work itself is larger than the person answering has in mind—and the two
compound rather than compete. A four-hour job that is really a nine-hour job,
done in a day full of interruptions, is not thirteen hours; it is nine hours of
work delivered across a stretch of calendar nobody predicted, with the
resumption cost paid at every one of the boundaries.&lt;/p&gt;
&lt;p&gt;So the second reason technical estimates come in short: the person answering has
one step in mind, and the work has several, each with its own limit.&lt;/p&gt;
&lt;p&gt;Restoring from backup is the clearest case. The details depend entirely on the
backup method, but the candidates for the constraint include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Read speed from the device holding the backup&lt;/strong&gt;, which is a property of
  both that device and the backup server's own I/O system.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The backup server itself&lt;/strong&gt;, which has to do the work of serving the
  restore.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Network bandwidth&lt;/strong&gt; between the backup server and the system being
  restored.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Write speed on the system being restored&lt;/strong&gt;, which is frequently the limit
  nobody thought about, because the estimate was framed as a question about
  backups.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Whichever of those is slowest sets the time, and the person giving the estimate
is rarely thinking about all four.&lt;/p&gt;
&lt;p&gt;Off-site backup copies complicate this in a way worth separating out. Holding a
copy somewhere else is good practice and the right answer for disaster
recovery, since a backup in the same building as the original is not a backup
against several of the things most likely to destroy both. It also frequently
means the restore path runs over a much narrower link than the local one. An
organization can be entirely correct to keep the off-site copy and still be
badly wrong about how long restoring from it takes, and the estimate given in
the room is almost always the local number.&lt;/p&gt;
&lt;p&gt;Few organizations have ever calculated their real maximum, and it is never the
number on the manufacturer's data sheet. That figure describes ideal conditions
no production system reproduces.&lt;/p&gt;
&lt;p&gt;It can be done, though, and the exercise finds the people who have done it. In
a recent exercise the head of the server team was a participant. Asked about
restore time, he said plainly that he could not speak for the network bandwidth
between the systems—and then said that he had calculated the I/O bandwidth
himself when specifying the backup server. That is exactly the right answer: a
real number for the part he owns, and an explicit boundary where his knowledge
stops. It is also uncommon.&lt;/p&gt;
&lt;p&gt;Then there is a compounding effect, in the class of incident where it applies.
Where the response requires restoring many systems at once—ransomware and
destructive malware are the obvious cases—simultaneous restores contend for the
same shared resources: the backup server, the storage behind it, the network
between them. Depending on the hardware and the topology, that contention can
be considerably worse than proportional, and an estimate derived from restoring
one system will not survive it.&lt;/p&gt;
&lt;h2&gt;Hardware that no longer exists&lt;/h2&gt;
&lt;p&gt;"How long to get the replacement?" deserves asking on its own, because the
answer has moved.&lt;/p&gt;
&lt;p&gt;The first question is whether cold spares exist. If they do, the estimate is
about installation. If they do not, the estimate is about somebody else's
supply chain, which is a different kind of number entirely: one the
organization does not control and probably has not checked recently.&lt;/p&gt;
&lt;p&gt;It has been a bad few years to be checking late. The build-out of AI data
centers has pulled memory and other components toward buyers who order them by
the container, and lead times on ordinary enterprise hardware have stretched
accordingly. An organization planning a routine server refresh now queues
behind hyperscalers for the same parts.&lt;/p&gt;
&lt;p&gt;The second question is subtler, and it catches the people who did check.
&lt;strong&gt;Hardware model lifespans are short&lt;/strong&gt;, often measured in months and rarely
much beyond a year. The exact model in the rack may simply not be purchasable,
whatever the lead time. The replacement task is then not procurement but
engineering: working out what will function correctly in this situation, with
this software, in this rack, on this power budget—and doing it under incident
conditions, which is the worst circumstance in which to make a design decision.&lt;/p&gt;
&lt;h2&gt;The estimate also assumes skill that may not be there&lt;/h2&gt;
&lt;p&gt;Searching logs is a good second example, because the constraint is not only
hardware.&lt;/p&gt;
&lt;p&gt;Suppose the organization has a SIEM and the logs are in it. Somebody is asked
to establish the initial entry vector. That is not a narrow lookup: it means
searching across gigabytes, often considerably more, over a period nobody has
bounded yet.&lt;/p&gt;
&lt;p&gt;Two things then decide the answer, and neither is in anyone's estimate. The
first is how well the SIEM indexes the data being searched, which varies
enormously between products—and the ones that hold up are expensive and need
capable hardware underneath them, so this is a purchasing decision made years
earlier that presents its bill during an incident.&lt;/p&gt;
&lt;p&gt;The second is whether the person searching knows the query that returns what is
needed and nothing else. If they do, it is minutes. If they do not, it is a
broad query, thousands of returned entries, and a person reading them, which is
not the same task at all. The tooling is identical in both cases. The
difference is whether a particular skill happens to be present at that moment,
on that shift.&lt;/p&gt;
&lt;h2&gt;Dependencies on external people or organizations&lt;/h2&gt;
&lt;p&gt;Any step that depends on somebody outside the team carries a hidden assumption:
that they will respond immediately, and that they will be competent when they
do.&lt;/p&gt;
&lt;p&gt;Where the dependency is a service provider, the assumption is usually
contradicted by the contract. A managed service provider has an agreed maximum
response time, and it is entirely proper for them to use it—a response inside
the committed window has met the agreement. It is also unlikely that the first
response contains everything needed, so the real elapsed time is several
exchanges long, each governed by the same clock. Internal groups frequently
have service levels to other internal groups, with the same effect and less
visibility.&lt;/p&gt;
&lt;p&gt;Then there is the question of who is even at work. A recent exercise reached
the question of which contracts governed the data held on a particular system.
That is not an IT question, and the people who can answer it keep office hours.&lt;/p&gt;
&lt;p&gt;How much that costs varies more than any other delay in this article, and the
range is worth stating rather than assuming the worst. Some organizations can
raise somebody in Contracts at ten at night—a small company where everyone has
everyone's mobile number, or one that has thought about which non-IT roles an
incident needs and put them on a call list. Others cannot, and the answer waits
for the morning. The figure offered in that room was an hour, which assumed
reaching the right person at once and getting a complete answer first time; the
useful correction is not a bigger number but the two questions underneath it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Can this person be reached out of hours at all, and has anyone checked?&lt;/strong&gt; A
name on a plan is not a phone number that gets answered, and the reasons it
might not be answered have changed in a way most call trees have not caught up
with.&lt;/p&gt;
&lt;p&gt;Phones now silence themselves. A sleep schedule or a focus mode turns the ringer
off every night automatically, and the exception list is a handful of close
contacts—family, not the on-call rotation. Adding somebody to that list is a
deliberate act that nobody does for a colleague they might need once. And even
where the exception exists, the phone is frequently charging in another room,
which defeats it completely.&lt;/p&gt;
&lt;p&gt;None of that is carelessness. It is the default configuration of an ordinary
phone belonging to somebody with a healthy relationship to their work. The
consequence is that "we can call him" needs testing rather than assuming: has
anyone actually rung that number at eleven at night and seen what happens? An
organization that has tested it knows whether it has an out-of-hours path. One
that has not is holding a list of numbers and a hope. And &lt;strong&gt;is the first
answer likely to be the complete one?&lt;/strong&gt; A contracts question rarely is: the
document has to be found, the relevant clause read, and its application to this
specific system decided, which is frequently a second conversation with somebody
else. Even the organization that can reach its Contracts lead at ten at night is
usually looking at more than an hour before it has an answer it would act on.&lt;/p&gt;
&lt;h2&gt;The number that describes work nobody has done&lt;/h2&gt;
&lt;p&gt;Some estimates describe a capability demonstrated only under conditions unlike
the ones that will apply.&lt;/p&gt;
&lt;p&gt;Kenneth spoke with someone able to image a remote system's storage and capture
a memory image remotely. A genuinely useful capability, and the person is
competent to do it. He has not, however, done it against a remote system: the
timing he would quote comes from doing it across a local network.&lt;/p&gt;
&lt;p&gt;That is not an idle distinction. The whole point of the capability is reaching
something that is not on the local network, and the link is exactly the
variable that has never been exercised. &lt;strong&gt;The estimate is honest, from a
skilled person, about work they can genuinely do—and it describes a different
job from the one that will be asked of them.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;The question is a probe, not a measurement&lt;/h2&gt;
&lt;p&gt;The estimate is not really the point. What it is made of is the point, and the
follow-up questions are where the exercise earns its day:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Has anyone here done this at this size?&lt;/li&gt;
&lt;li&gt;When was the last time?&lt;/li&gt;
&lt;li&gt;What has changed since?&lt;/li&gt;
&lt;li&gt;Which part of it is the slow part, and how do you know?&lt;/li&gt;
&lt;li&gt;Who else has to be available for it, and what happens if they are not?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;What the exercise cannot do is produce the true number.&lt;/strong&gt; There is no time in
an exercise to run a restore or a search and measure it, and a room that spends
its afternoon trying has stopped doing the exercise. That is a real limit and
it should be said plainly rather than glossed over.&lt;/p&gt;
&lt;p&gt;What the exercise can establish is whether a defensible number exists anywhere.
A room that can answer the questions above has one, or knows where it is
written down. A room that cannot has discovered something more useful than a
number: that its recovery planning rests on a figure nobody has tested, which
is a finding, and one nobody had to run a stopwatch to reach.&lt;/p&gt;
&lt;p&gt;The second useful move is asking two people separately. The backup
administrator and the person who would actually be woken up frequently give
different answers, and the difference is rarely a disagreement about
technology. It is usually that one of them is including the conditions and the
other is not.&lt;/p&gt;
&lt;h2&gt;The after-action item writes itself&lt;/h2&gt;
&lt;p&gt;This is one of the rare exercise findings with a clean, cheap recommendation
attached, and the recommendation is not "revise the estimate". It is &lt;strong&gt;do the
thing and time it&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Pick the steps the plan depends on, carry them out under conditions as close to
real as can be arranged, and record what they actually took. The timing is only
half the return. The other half is confirming that the expected data is really
there: that the backup contains what everyone believes it contains, that the
logs retain the period that would need to be searched, that the remote capture
works against something remote.&lt;/p&gt;
&lt;p&gt;Where an IR plan states time estimates for its steps—and plans increasingly
do—those numbers need real tests behind them, and &lt;strong&gt;the conditions of the test
need recording somewhere alongside the estimate&lt;/strong&gt;. An unsupported number in a
plan is worse than no number, because it will be relied on by people who assume
somebody measured it. "Forty minutes" means something quite different depending
on whether it was one system or twelve, a local network or a remote one, a
practiced operator or whoever was on shift.&lt;/p&gt;
&lt;h2&gt;A note on how to ask&lt;/h2&gt;
&lt;p&gt;Ask for the number before discussing the difficulties, not after. Once the room
has spent five minutes on everything that could go wrong, the estimate that
follows is contaminated by the discussion, and the interesting artifact—the
number people carry around in their heads and plan with—is gone.&lt;/p&gt;
&lt;p&gt;Ask, write it down, then open the discussion. The first answer is the one the
organization has actually been relying on, and the distance between it and the
second answer is the finding.&lt;/p&gt;
&lt;p&gt;Ultimate TTX runs &lt;a href="/facilitated-booking/"&gt;facilitated exercises&lt;/a&gt; built to
surface exactly this kind of gap, on the day, in front of the people who can
act on it.&lt;/p&gt;</content><category term="Facilitation"/><category term="estimates"/><category term="recovery time"/><category term="injects"/><category term="findings"/></entry></feed>