Site readiness requires evidence for the actual configuration. An empty rack position does not by itself mean that a server can be installed and loaded. Use this process as a shared work order for the infrastructure owner, site engineer and supplier, keeping decisions in one readiness record.
1. Check the rack and delivery route
Start with the exact chassis model and final bill of materials. Request installation drawings and requirements for rails, dimensions and maintenance access. Compare them with the actual rack and the ability to remove the unit safely for service. Assign someone to sign off the location before shipment.
Record the route from unloading to the rack, doorway and lifting-equipment limits, installer access and packaging space. Agree unloading and installation separately; do not assume either is included. NVIDIA’s DGX guide addresses rack stability, loading and grounding; use the documentation for your chosen chassis.
2. Agree power and cooling
There is no universal power figure for a GPU server. Ask the site engineer to approve feeds, PDUs, cables, connectors, protection and redundancy against the node specification. Record the measurements required at acceptance and who authorises a load test. Adding GPU nameplate values does not replace electrical design.
Request a separate cooling sign-off. For an air-cooled system, discuss air supply and exhaust; for a liquid-cooled system, define responsibility for the loop and its maintenance. The table lists questions for the delivery team, not a universal connection design. Work must follow equipment and site documentation.
| Check | Air cooling | Liquid cooling |
|---|---|---|
| Starting information | Chassis requirements and inlet-air conditions | Manufacturer requirements for the loop and coolant |
| Infrastructure | Air delivery and rack heat removal | An agreed interface with the site cooling loop |
| Delivery scope | Ducts, blanking panels and cable organisation | Connection assemblies, monitoring and accessories |
| Acceptance | Temperatures and behaviour under the intended load | Agreed loop checks and workload tests |
| Responsibility | Who resolves cooling issues | Who maintains the server and site portions |
3. Separate data networking from management
Draw the connections for compute traffic, storage, user requests and management. Record the port, cable, transceiver, purpose and owner for each connection. Agree addressing, routes and storage access beforehand; a network adapter does not prove that the complete data path is ready.
Document BMC and operating-system access separately. NVIDIA recommends a dedicated management network protected by a firewall and secure remote BMC access, such as through a VPN. Define account owners, permission grants and revocation, emergency access and alert recipients. Test access from the administrator’s real network.
4. Record a reproducible software environment
Before testing, inventory server and GPU firmware, BIOS/BMC, OS, driver, CUDA, ROCm or CANN, container environment and framework versions. Ask the implementer to confirm that the combination is compatible, rather than simply installing the latest release of each component. Record the exact model, weight format and application launch settings.
Save configuration, image checksums and a repeatable deployment procedure. Assign an update owner and agree how to return to a working state. Verify backups by restoring a test dataset. The acceptance record should contain commands or steps that reproduce the result after handover.
5. Accept the outcome, not merely a powered-on server
Separate acceptance into configuration compliance, node diagnostics and your application test. NVIDIA describes preliminary health checks and a stress test for DGX; these do not replace an application benchmark. Agree duration, metrics, tolerances and retest conditions before running the tests.
Keep serial numbers, logs, versions and measurement results. Verify restart behaviour, restored access and alert delivery. Clarify warranty coverage at the operating location: who receives a case, what evidence is required, how repairs are arranged and who owns data responsibilities. Response time, recovery time and replacement equipment are distinct terms; record each separately.
A site is ready when responsibilities are assigned and the agreed checks pass reproducibly.
Six confirmations before launch
Sources and documentation
- NVIDIA DGX: safe site installation requirements ↗
- NVIDIA DGX: securing management access ↗
- NVIDIA DGX: diagnostics and pre-flight testing ↗
These guides help you prepare requirements. The exact configuration and terms are set out in the quotation.
