The Operational Question

When a data center loses a utility source, cooling path, controls network, building access system, or another facility capability, can the team use its disaster recovery runbook without improvising the first critical decisions? A document that looks complete in a shared drive may still omit the actual switch locations, operating limits, communications handoffs, vendor contacts, or recovery evidence that technicians need during a real event. Testing the runbook turns a business-continuity document into an operating tool for the physical plant.

This guide explains how facilities managers, operations leaders, and training coordinators can test a disaster recovery runbook without creating an unnecessary live-system hazard. It covers the people who should participate, the difference between a discussion exercise and an equipment test, the evidence to collect, and the course choices that support a role-based plan. The goal is not to rehearse every possible catastrophe. The goal is to expose one specific recovery path that would otherwise fail under pressure.


Who This Affects

Runbook testing matters in every facility where physical infrastructure supports customer commitments, production systems, research, or public services. The exercise should be adapted to the site rather than copied from another building.

The participants should reflect the decision chain, not just the people who wrote the document. A facilities manager may authorize a recovery mode. A critical facilities technician may verify switchgear, generators, UPS modules, chillers, or CRAH units. A NOC operator may receive the first alarm and create the incident bridge. Security may control access for responders. An EHS manager may stop an unsafe task. A vendor or electrician may be needed for a specialized repair, but the site still needs an internal owner for each decision.


What Can Go Wrong

The most common runbook failure is not a missing paragraph about a rare event. It is a mismatch between the written sequence and the facility as it is operated today. Equipment may have been replaced, labels may have changed, a bypass may be locked, or a contact list may point to a person who no longer supports the site.

During a utility-loss scenario, a team may call for generator operation without confirming fuel status, automatic transfer behavior, load-shed priorities, exhaust conditions, or the authority to enter the generator yard. During a cooling incident, the runbook may say to move load to standby capacity without identifying the actual valves, control points, alarm thresholds, or temperature limits that make the move safe. During a controls-network failure, operators may rely on a dashboard that is unavailable and discover that local gauges, manual controls, and rounds were never assigned.

These gaps create several kinds of exposure:

A discussion exercise also has limits. Talking through a generator start does not prove that the generator will start. Reviewing an EPO response does not authorize a live EPO test. Confirming that a clean-agent alarm appears on a screen does not demonstrate that the room, notification, evacuation, and incident command procedures are ready. Runbook testing should state what is being simulated, what is being observed, and what physical testing requires a separate approved procedure.

Standards and guidance can inform the exercise, but the training package is not a regulatory certification or a substitute for site-specific qualification. OSHA requirements, electrical safety practices, emergency-management guidance, manufacturer instructions, contracts, and local procedures still apply to the work.


What Managers Should Check

Use a controlled, evidence-based exercise. Start with a narrow scenario, such as a utility interruption during a maintenance window, loss of a chilled-water pump, failure of a BMS server, or restricted access to a mechanical room. Avoid combining five failures in the first exercise. Complexity can be added after the team can execute the basic recovery path.

1. Set the exercise boundary

Write a short scenario statement with a start time, affected asset, known alarms, current operating mode, and constraints. State whether the activity is a tabletop discussion, a communications drill, a walkdown, a simulated alarm injection, or a separately approved equipment test. Include a hard stop that prevents a participant from operating equipment unless the existing procedure and authorization permit it.

The exercise owner should identify the success condition. For example, the objective might be to establish a safe cooling response within ten minutes, account for affected personnel, communicate customer impact, and document the decision to hold or change the operating mode. It does not need to be “restore everything immediately.” A good objective makes safe delay visible when the evidence is incomplete.

2. Validate the starting information

Before the exercise, compare the runbook with the latest site information. Check the following:

If the runbook says “check the panel,” the team should be able to identify which panel and where it is. If it says “verify redundancy,” the team should be able to name the remaining path and the evidence that proves it is available.

3. Assign roles before the clock starts

Use named roles, not a vague instruction to “notify the team.” Assign an incident commander, facilities lead, safety lead, communications lead, scribe, security coordinator, and technical subject-matter contacts as needed. One person can hold more than one role at a small site, but the exercise should make that tradeoff explicit.

The scribe should record the time of each decision, the information available, the person who made the decision, the action owner, and the next verification. This creates a useful after-action record without asking technicians to write a narrative while troubleshooting.

4. Walk the physical path

A tabletop should be followed by a controlled walkdown. Start at the alarm or notification point and trace the route to the equipment, isolation point, alternate path, staging area, and exit. Look for locked doors, unclear labels, blocked access, poor lighting, incompatible radios, missing drawings, and equipment that cannot be viewed safely from the expected position.

The walkdown should also test the handoff between control room and field staff. The operator may see a “pump failed” alarm, while the technician needs the pump number, operating state, valve lineup, local control mode, and permission to enter the plant. The exercise is successful when those details move through the team without being invented on the spot.

5. Test communications and decision points

A recovery runbook should identify what must be communicated, to whom, by when, and through which channel. Test the primary channel and the fallback. A phone tree may be current while a radio channel is not. A vendor may acknowledge a call but still lack the site history needed to respond.

Ask decision questions that reveal judgment, not memorization:

  1. What fact would make you stop the recovery sequence?
  2. Who can authorize a change to the operating mode?
  3. What customer or internal notification is required before the next step?
  4. What equipment state must be verified locally rather than assumed from a dashboard?
  5. What evidence proves the site is stable enough to return to normal operation?

The answers should be specific to the facility. “Follow the SOP” is not enough if the SOP does not name the current equipment or authority.

6. Capture recovery evidence

Define the records that demonstrate completion. They may include alarm screenshots, operator logs, equipment readings, a communication timeline, inspection notes, contractor tickets, fuel or temperature records, and a list of open impairments. Do not collect more information than the review team can use. A small set of reliable evidence is better than a large folder of unverified screenshots.

The return-to-normal step needs the same discipline as the initial response. Identify who confirms stable load, normal cooling, correct breaker or valve lineup, cleared alarms, restored monitoring, access control, and customer communications. If the facility has a temporary configuration, the runbook should show who owns the removal and when it must be completed.

7. Turn observations into training assignments

After the exercise, separate document defects, equipment defects, staffing gaps, and knowledge gaps. An outdated drawing requires document control. A failed alarm requires technical investigation. A missing spare requires maintenance or procurement action. A technician who cannot explain a switching or cooling sequence may need training, supervised practice, or a qualification review.

Give each finding an owner, due date, risk statement, and verification method. A revised runbook is not closed until the team uses the revised step in another discussion, walkdown, or approved test.


Which Training Fits This Situation

For an operations leader building a facility-focused recovery plan, Disaster Recovery & Business Continuity provides the broadest foundation for linking critical functions, recovery priorities, communications, and continuity decisions. Business Continuity and Disaster Recovery Planning Training is a useful modular option when the immediate need is a focused planning skill for coordinators, supervisors, or cross-functional participants.

The physical recovery path should determine the supporting assignments. A facilities manager may pair the planning course with Data Center Operations Management to connect the exercise to operating modes, logs, escalation, and service commitments. A technician assigned to verify electrical states may need Power Distribution Systems Fundamentals or Emergency/Standby Power Systems Fundamentals. A cooling lead may benefit from Cooling Systems Design & Optimization or HVAC Systems Troubleshooting Essentials. A controls specialist may need Monitoring, Automation & BMS Systems or DCIM Platform Fundamentals.

When the site is building a common baseline across operations, facilities, and support staff, the Operations & Reliability Bundle is the most direct bundle-level fit. It can support a role-based plan, but enrollment does not by itself qualify anyone to operate a particular switchgear lineup, generator, chiller, or control system. The employer still needs site orientation, manufacturer procedures, supervised practice, and its own authorization process.

Use a simple matrix to assign training:

Learners receive a certificate of completion for the selected online course. That document records training completion; it is not a license, regulatory certification, or endorsement by OSHA, NFPA, NIST, or another standards body.


Common Mistakes to Avoid


Key Takeaway

A disaster recovery runbook becomes useful when a real facilities team can follow it, challenge it, and show evidence for each important decision. Start with one narrow failure scenario, walk the physical path, test the communication handoffs, and assign every gap to documentation, equipment, staffing, or training. This week, schedule a 30-minute tabletop for one critical facility dependency and invite the person who would have to verify the recovery in the field.


Sources