SLA Evidence Chain
Remote Diagnosis Classifies Risk; On-Site Service Isolates, Tests and Verifies Recovery
A data center fault can first be triaged remotely through alarms, trends and impact scope. But electrical safety, ATS/STS, paralleling, breakers, cooling, fuel leaks or compromised redundancy require an on-site process. The boundary must protect the SLA and safe isolation.
Probable Causes
Problems Affecting IT, Cooling, Redundancy and Isolation Cannot Rely on Remote Reset
Remote data can quickly classify trends and risk, but breaker operation, distribution testing, paralleling adjustment, cooling controls and safe isolation all require on-site confirmation. A data center cannot sacrifice redundancy or operating safety for a quick recovery.
Remote Support Suits Trends and Risk Classification
Operating parameters, alarms, transfer records, fuel level and temperature trends help define the impact scope.
On-Site Service Suits High-Risk Operations
Low-voltage switchboards, ATS/STS, breakers, paralleling switchgear, fuel and cooling systems require on-site instruments and safety procedures.
Redundancy Must Be Verified After Recovery
After the issue is handled, confirm that IT loads, cooling and redundant paths again meet SLA requirements.
Diagnostic Process
Engineers Assess SLA Impact Before Choosing Remote or On-Site Action
- 01
Confirm whether the fault affects IT loads, cooling, redundancy, fuel runtime or safe isolation.
- 02
Remotely review alarms, trends, transfer records, load ratio, fuel level and temperature.
- 03
Determine whether on-site instrument testing, breaker operation, paralleling adjustment or isolation measures are required.
- 04
After on-site work, verify recovery of IT loads, cooling, distribution and redundant paths.
When an On-Site Inspection Is Required
These Conditions Require On-Site Inspection, Not Remote Observation Alone
- The fault affects IT loads, cooling, N+1 redundancy, ATS/STS or low-voltage distribution.
- Breaker operation, paralleling adjustment, fuel-system inspection or safe isolation is required.
- Remote data is insufficient to prove the fault section and recovery result.
Prevention and Maintenance
To Find Problems Before an Outage, Every Inspection Must Be Recorded and Reviewable
Preventive Checks
- Predefine fault-classification rules and identify conditions that require on-site intervention.
- Keep remote-monitoring data complete to support rapid classification and on-site preparation.
- Regularly rehearse remote diagnosis, site access, safe isolation and recovery-verification processes.
Maintenance Routine
- After a fault, record the impact scope and redundancy status before performing any reset.
- Before on-site intervention, define operating boundaries, rollback plans and communication mechanisms.
- After service, create recovery-verification records proving that the SLA risk is closed.
Engineering Recommendation
Remote reset is not a universal solution for data centers. Remote support diagnoses and on-site service closes the loop; together they prevent a minor fault from becoming a redundancy incident.