Est.

Operational Readiness Review Process for Critical Facilities

Readiness means systems, people, and data actually work together on day one.

Columnist · · 11 min read
Cover illustration for “Operational Readiness Review Process for Critical Facilities”
Operations-Ready Handoff · September 11, 2026 · 11 min read · 2,375 words

An Operational Readiness Review isn't a checklist you run through the week before opening day. It's the process that proves, from the first design review through the day operations takes the keys, that the people, the systems, and the data behind them actually connect and trace back to a source. That's different from "mechanically complete," and the gap between those two milestones is where post-occupancy outages and expensive callbacks come from.

Readiness, in the sense that matters to an owner, means the facility can run the way the Owner's Project Requirements say it should, right when construction hands the keys to operations. Facility Grid's framework splits that into three parts that lean on each other: systems and equipment, people, and asset data. None of the three stand alone. Salute frames operational readiness the same way from a different angle: aligning people, process, technology, and governance, not just keeping the lights on. None of this is new to critical facilities, either. A 2024 IEEE paper by Cusick and Basil traces ORR back through IT systems and cloud release management, long before it became standard language on data center construction sites.

Why commissioning stages are the structural backbone of an ORR, not a parallel track

ASHRAE lays out data center commissioning as a sequence of levels, each one building on the last. The ORR doesn't run alongside that sequence. It sits on top of it, enforcing the gates.

Facility Grid's expansion of the ASHRAE structure moves through Level 0 (submittal reviews, OPR, design reviews, all before construction starts), Level 1 (factory acceptance and bench testing), Level 2A (delivery and installation verification), Level 2B (quality control and inspection), then completion milestones at Levels 5 and 6. One rule holds the whole thing together: each level has to finish before the next one opens. No skipping ahead, no "we'll circle back." An ORR that lets a team defer completion at one level and pick it up later isn't doing its job. It's just writing down the shortcut for someone to find later.

What commissioning actually checks, at every level, is whether cooling, power distribution, and critical equipment work together under real conditions, not whether each piece passes its own test sitting by itself on a bench. A UPS that clears a bench test and a chiller that clears its own commissioning script can still fail each other the moment they're asked to perform together during an actual utility event.

The scale is worth sitting with. There are 10,360 data centers worldwide, 5,381 of them in the U.S., and McKinsey projects 10% annual growth through 2030, backed by $49 billion in new construction spend. Every one of those projects runs commissioning gates. A failed gate on a single project used to be a local headache. Now it's a failure mode that repeats across a build pipeline that keeps getting bigger and faster, with new cooling tech and tighter system integration raising the stakes at each step instead of lowering them.

How fragmented tracking tools undermine the ORR process before it reaches the finish line

Ask anyone running commissioning on a data center project what tool tracks progress, and the honest answer, most of the time, is still a spreadsheet. And a spreadsheet goes stale the second someone clicks save. On an active job site, tests finish, checklists update, and issues get logged faster than a static file can follow, so the version a project executive is staring at rarely matches what's happening on the floor right now.

That gap is where surprises live: rework nobody budgeted for, delays nobody scheduled, penalties nobody wanted to explain to the owner. Those are exactly the outcomes an ORR exists to catch before they happen, not after.

The fragmentation runs deeper than tracking. The same piece of equipment data often gets rebuilt by hand across a model, a spreadsheet, a PDF, and whatever operational system is supposed to receive it eventually. A single rack move, something that sounds trivial, can ripple into rework across power, cooling, cable routing, and documentation, none of it updating on its own. RFI answers and design decisions that live in disconnected files go untraceable within months, sometimes weeks.

The MEP coordination conflicts that surface during construction, instead of getting caught during design, are a direct symptom of that fragmentation. One documented 10 MW data center project ran a chilled water pipe straight through the primary cable tray path serving 40% of the server cabinets. Fixing it cost $380,000 and pushed IT equipment installation back five weeks. Coordinated MEP modeling has been shown to cut field RFIs by 30 to 50 percent on complex projects, and in a mission-critical environment, every RFI avoided is a schedule threat avoided. An ORR can't validate information that project management already let fall apart. Readiness of the information has to come before readiness of the systems and the people using it, not the other way around.

What a complete ORR actually covers: power, cooling, cable, and the systems that depend on each other

Power validation starts with confirming A/B path separation end to end, not just checking it at the switchgear and calling it done. Protection system tuning belongs in the ORR itself, not buried as a commissioning footnote, because a poorly tuned protection scheme risks tripping load at scale during a fault. A 2024 event in Virginia, reported by NERC, saw 1.5 GW of data center load disconnect within 82 seconds of a single 230 kV line fault, which triggered successive system faults through auto-reclosing. UPS and generator protection settings need to match the utility's reclosing and ride-through curves, or that kind of cascading disconnect turns from a rare event into a repeatable risk.

Cooling validation has to treat CRAH units, chilled water piping, and rear-door or direct-to-chip liquid cooling as one system, because that's how they actually behave together. Thermal hotspots show up when airflow meets an obstruction nobody flagged, and while that surfaces during commissioning, the root cause almost always traces back to a design coordination gap that could have been caught earlier. AI workload density has changed what "cooling validated" even means: hybrid capability, liquid and air together, has to be checked on its own terms. A general commissioning sign-off doesn't cover it.

Cable and fiber routing gets its own hard numbers. Power trays and fiber or copper trays need minimum clearance between them, which roughly doubles the pathway space required compared to running everything in one unified tray. Trays shouldn't exceed 40 to 50% fill at installation, so the installed capacity needs to run 2 to 2.5 times the initial cable volume, a requirement design teams underestimate more often than not, and one that tends to surface right at ORR. Separation, access for service, bend radius, connection rules: sign-off on all of it needs to land before IT equipment ever touches the floor.

All of it sits inside a single ORR. And the load these systems carry keeps climbing. Goldman Sachs Research projected, in February 2024, that AI would drive a 165% increase in data center power demand by 2030. Whatever gets validated at ORR today has to serve a load denser and more dynamic than what the original design ever assumed.

Documentation completeness as an ORR gate: what "done" means for asset data before turnover

The three readiness domains only work together. Operations staff can't run equipment they can't find records for, and DCIM, EPMS, and BMS can't manage assets nobody entered into the system in the first place.

Documentation completeness at ORR means checking every piece of MEP equipment against as-built conditions, not the original design model, because the two rarely match exactly by the time construction wraps. Equipment parameters, make, model, capacity, firmware version, protection settings, need to trace back to the submittals and RFI answers that decided them. O&M documentation has to be received, organized, and linked to the specific asset it describes, not dumped into a folder of PDFs someone can dig through later if they're patient enough. Cable schedules and port assignments need to show what actually got installed in the field, not what the drawing said would happen.

Equipment specs trapped inside manufacturer PDFs are wasted information at this stage. The goal is pulling that data out so it flows straight into models, schedules, and validation rules, because handing operations staff a stack of PDFs isn't a handoff. It's shifting manual labor from one team to another. BIM data that reaches turnover incomplete, or sitting disconnected from operations, isn't BIM doing its job. The finish line is a structured handoff into DCIM, EPMS, and BMS, not a 3D model parked in a shared drive nobody opens again.

Good BIM practice in data centers means building cabling and patching scopes directly into the model, running CFD simulations to check the cooling strategy actually holds up, and making sure the final model reflects what got built, not what got drawn eighteen months earlier. Research into BIM maturity lands on the same point from another angle: teams working off one shared source of truth are better positioned to capture information in a form that supports operational handoff, rather than requiring manual re-entry. Every RFI answer, every design decision, every equipment parameter should trace back to whatever backs it up. The ORR is the structured point to check that traceability exists, before someone needs it in the middle of an incident and can't find it.

Integrating into DCIM, EPMS, and BMS: what the ORR must confirm before operational systems go live

Three systems, three different jobs. BMS handles facility-level HVAC, lighting, fire, security, and environmental controls, but it doesn't touch IT-layer operations. EPMS covers electrical distribution and energy metering, built for power chain telemetry that BMS was never designed to capture. DCIM sits between the two, pulling together power chain data, cooling and environmental readings, and asset lifecycle information across sites, bridging the facility layer and IT service management.

A common mistake at ORR is assuming DCIM replaces BMS, or replaces IT service management. It replaces neither. ORR validation needs to confirm the three systems actually talk to each other, instead of taking a vendor's single-system claim at face value.

That means testing the data collection protocols each system leans on, covering the full range of interfaces used by installed equipment. Every collector interface needs testing against the equipment that's actually installed, not assumed to work because the design spec said it would. AI-driven workload density has made this harder. Platforms built around the old assumption of static, air-cooled racks handle the shift toward direct-to-chip liquid cooling, rear-door heat exchangers, and immersion cooling unevenly: some well, some barely at all. The ORR needs to check hybrid cooling telemetry specifically, not accept a general "AI-ready" claim off a vendor data sheet.

Industry coverage of DCIM platforms points to differing strengths across vendors, with some oriented toward enterprise monitoring, others toward workflow automation, and others toward native EPMS integration. Full DCIM deployments can run three to six months per site, so integration planning at ORR isn't a same-week decision. It's a scheduling constraint that shapes the whole turnover timeline.

Firms that run the same BMS and EPMS architecture across multiple sites, whether new builds, retrofits, or live conversions, carry less integration debt into each ORR. That consistency cuts the variability that trips teams up when every site turns into a slightly different puzzle. Industry observers have noted that DCIM's role has mostly stayed limited to real-time monitoring, asset management, and capacity planning. The ORR is the right moment to ask whether the platform actually does more than watch, and whether it supports operational decisions from day one rather than just logging them.

The Open Compute Project's OCP Ready program adds an outside layer on top of internal ORR work, formalizing facility readiness for colocation providers serving hyperscale deployments. It runs two tiers: v1 for general OCP colocation, v2 for hyperscale-specific facilities, the more rigorous of the two. Getting recognized under the program means review by the OCP Data Center Facility Project community and sign-off from the OCP Ready Facilities Lead, a Steering Committee Representative, and an OCP Foundation Representative. That's a genuinely separate check, not a rubber stamp sitting on top of internal review.

Where automation changes the economics of running an ORR on a compressed schedule

Compressed schedules aren't a phase the industry is passing through. They're the environment projects get delivered in now. The challenge is plain: with labor shortages and timelines only getting tighter, teams need something better than what they've been using to run commissioning and quality control.

Static tools fail here because commissioning itself isn't static. Tests finish, checklists update, issues get logged, all faster than a spreadsheet can keep pace with. Real-time visibility across systems, equipment, and documentation isn't a nice feature to have someday. It's what the schedule itself demands.

Automation earns its place in an ORR when it does four things. It cascades design changes, so a rack move updates power, cooling, cable, and documentation records on its own instead of triggering manual rework across four separate disciplines. It enforces gate sequencing, so each commissioning level shows verifiably complete status to project executives and schedulers in real time, not two weeks after the fact. It keeps traceability intact, linking RFI answers, submittal decisions, and equipment parameters to the assets they touch, so an ORR team can audit anything without digging through a folder structure three different people built over two years. And it backs a structured handoff, letting asset data flow from the design environment straight into DCIM, EPMS, and BMS at turnover, instead of getting typed in by hand a second time.

AI and deterministic automation aren't the same tool doing the same job, and treating them as interchangeable is where teams get into trouble. AI earns its spot where flexibility and judgment matter: flagging anomalies, surfacing patterns buried in commissioning data that a human might miss on a first pass. Deterministic rules belong where precision and safety aren't up for interpretation: power path separation, protection settings, tray fill ratios, clearance requirements. Mix the two up, and let a probabilistic tool make a call that needs to be exact every single time, and the ORR ends up creating risk in exactly the spot it was built to remove it.

Sources

  1. Data Center Facility/OCP Ready - OpenCompute
  2. facilitygrid.com
  3. Discover the top 10 best practices for operational readiness
  4. A Global Operational Readiness Review Process: Improving Cloud Availability
  5. facilitygrid.com
  6. facilitygrid.com
  7. datacenterfrontier.com
  8. modius.com

More in Operations-Ready Handoff