As liquid cooling technology scales across data centers, the critical importance of pre-commissioning cleanliness has become increasingly apparent to industry leaders. Construction debris and contaminants inadvertently introduced during system installation can wreak havoc on narrow-channel cold plates, leading to blocked channels, reduced flow rates, and elevated temperatures that impact compute performance from day one. Bob Walicki, RD&E Program Leader at Ecolab, has been working directly with customers and observing commissioning challenges firsthand, contributing to the development of comprehensive OCP white paper guidelines that establish repeatable procedures from hydrotest through final fill.
The stakes of getting pre-commissioning right extend far beyond initial performance issues. Poor practices during this critical phase can create environments conducive to microbial growth and ongoing contamination accumulation, ultimately requiring intensive maintenance interventions and costly unplanned downtime. From staged filtration strategies and real-time coolant telemetry monitoring to the role of predictive analytics in optimizing long-term fluid health, Walicki shares practical insights for engineering managers planning their first direct-to-chip deployments, emphasizing that preventing a single rack rework often offsets the entire cost of proper pre-commissioning procedures.
++
VMblog: From a research standpoint, why has pre-commissioning cleanliness for row/rack cooling loops become such a critical focus as liquid cooling technology scales?
Bob Walicki: Through working with customers and observing commissioning of new systems, we’ve seen that construction/fabrication debris and other contaminants inadvertently introduced into the systems can pose challenges for the narrow-channel cold plates used to keep the compute technology cool. We’ve worked with partners to deliver our OCP white paper which codifies repeatable steps from hydrotest to clean, flush and prep for final fill to help mitigate risk as data center providers scale up their deployments.
VMblog: How do pre-commissioning practices impact day-one performance and long-term reliability?
Walicki: If poor practices are followed during pre-commissioning, infiltration of particles and air can impact day-one performance by blocking cold-plate channels, thereby reducing flow and increasing local temperatures resulting in high thermals on the compute side. Longer-term, the threat increases to include potential for creating environments where, for example, microbial growth can occur or inorganic contaminants continue to accumulate, which would require even more intensive maintenance and downtime to mitigate.
By following proper pre-commissioning procedures, you can avoid scenarios that lead to early rework and protect yourself from unplanned downtime and more impactful maintenance long-term.
VMblog: Why does the paper recommend staged filtration down to 5 micron and what are the trade-offs?
Walicki: Staged filtration is recommended as a best practice because immediately filtering to final operation levels, typically between 1 to 25 microns, is likely to result in rapid clogging and significant operational issues with pressure drop during the flushing step, as one can expect there will be some particle loading in the system. Starting with 50- to 100-micron filtration will enable efficient circulation of fluid through the system at a reasonable pressure drop. Systems may vary in their final filtration rating, usually between 1- and 25-micron, but in this white paper we recommend 5 micron. Considerations for specific systems involve balancing pressure drop, recirculation time, and how frequently the filter plugs and/or requires changeout.
VMblog: What are common hydrotesting mistakes that lead to later contamination or biofilm?
Walicki: The biggest mistake is not staging the subsequent steps properly and leaving water stagnant and not circulating in the loop for more than 24 hours. It is recommended to move directly to the recirculation in the cleaning step once hydrotest is passed.
VMblog: Where does real-time coolant telemetry (e.g., 3D TRASARTM technology) fit into the commissioning lifecycle? Commissioning tool, operations tool or both?
Walicki: Real-time coolant telemetry like 3D TRASARTM technology can be used to monitor and verify operations during hydrotest, flush, cleaning and fill. For example, monitoring pH and conductivity during rinse in real-time can help determine when all the cleaning chemistries have been removed from the system. Technology likes this also provides a record of coolant parameters during each of the steps to help with documentation.
When the system is commissioned, real-time coolant telemetry is an important indicator of system health and allows for proactive response to anomalies to eliminate damaging components leading to downtime and system outages.
VMblog: The OCP white paper is about pre-commissioning – how does continuous monitoring reduce operational risk afterward?
Walicki: Once you have taken the steps to properly commission a system, continuous monitoring replaces periodic sampling and provides real-time insights allowing you to connect shifts in coolant quality with operational actions and monitor trends that enable action before hardware is damaged or systems are impacted.
VMblog: For retrofits, which pre-commissioning steps are most overlooked? Recommendations?
Walicki: For retrofits, it is very possible to overlook some critical items, for example, drain and low points as well as dead-legs sometimes concealed in ceilings or below floor level. A comprehensive site survey can help identify all low points as well as dead-legs or other aspects that will impede full fill and drainage.
VMblog: How scalable are flushing/skid approaches for hyperscale vs small edge sites?
Walicki: It is possible to scale these approaches for any situation but resourcing in both cases can be challenging as timelines are always a challenge in this industry. For hyperscale environments, planning upfront will help determine the amount of parallelization possible with the amount of equipment that can be managed on-site. In smaller locations, it is likely that smaller skids and totes will suffice but there can be challenges that will require more intervention and oversight during the process.
VMblog: Can you quantify benefits of the following best practices (reduced failures, improved efficiency)?
Walicki: Following best practices during pre-commissioning helps reduce maintenance touchpoints in Year 1 and can help avoid shutdowns related to leaks, contamination or thermals related to cold-plate clogging. The exact return on investment will vary by site, but preventing a single rack rework or hardware or coolant replacement often offsets pre-commissioning costs.
VMblog: What is the role of predictive analytics and machine learning (ML) in reducing pre-commission variability and optimizing long-term fluid health?
Walicki: Using coolant telemetry combined with lab analysis and operational data to train ML models could help find unexpected or surprising relationships between the data being collected that could predict a variety of failure modes and communicate actionable advisories that will help prevent failures. As more data becomes available, these models could help improve current best practices and streamline pre-commissioning by identifying, for example, specific risks and their probabilities.
VMblog: What are three practical actions for an engineering manager planning a first direct-to-chip deployment?
Walicki:
- Plan, plan and plan. Create a written pre-commissioning plan that includes procedure water specifications, wastewater handling and a schedule that adheres to the recommended best practices.
- Book flushing skids and required consumables early; confirm instrumentation is on-site, operational and calibrated.
- Define telemetry points and sampling plan; install sample ports and plan for telemetry baseline data during initial fills.






