Education & healthcare
IT HelpdeskRestoring a degraded network under P1 major incident management
Client anonymised. Named references are available under NDA during due diligence.
The client delivers education and healthcare services in Australia across a distributed network that staff and clinicians depend on to reach systems and records. Canaan provides remote L1 and L2 IT helpdesk and infrastructure support across that estate — network operations, infrastructure, end-user computing, identity and access management, and ITSM — with network availability and business continuity as the standing brief.
The challenge
Users began reporting intermittent connectivity, application latency and service disruption across the network. Initial diagnostics pointed to high packet loss, latency spikes and bandwidth saturation, with access to business-critical applications degrading as a result.
The intermittent nature of the fault was the difficult part. A problem that comes and goes resists isolation, because the evidence disappears before anyone can look at it — and it resists it hardest in a remote support environment, where nobody can walk to the rack and watch a light blink.
The solution
Canaan activated P1 major incident management, with SLA-driven triage, impact assessment and structured escalation from L1 to L2 to L3 and vendor.
End-to-end network diagnostics — ping, traceroute, DNS and DHCP validation, and TCP/IP stack analysis — isolated the packet loss, latency and connectivity anomalies. LAN, WAN and bandwidth utilisation analysis identified the congestion, throughput degradation and infrastructure bottlenecks producing them.
Network telemetry was then correlated against server health metrics — CPU, RAM, disk I/O, latency and device availability — to narrow the fault domain faster than either data set could on its own.
Threshold-based proactive monitoring, performance baselining and automated alerting were implemented on the same conditions, so anomalous behaviour is detected before it becomes service degradation.
The outcome
Service was restored through the major incident process, and the fault domain was identified rather than worked around.
The more durable change is what the incident left behind. Threshold-based monitoring, performance baselines and automated alerting now watch the conditions that produced the degradation, which moves the next occurrence from something users report to something the desk is told about. An intermittent fault that was invisible until people complained became a measurable one.
Key takeaways
What carries across.
- Intermittent faults resist isolation precisely because they are intermittent — structure the triage rather than chase the symptom.
- Correlating network telemetry against server health narrows the fault domain faster than either data set does alone.
- Major incident management is a process, not a severity label: triage, impact assessment and defined escalation are what make it work.
- What the incident leaves behind matters more than the fix — the next occurrence should arrive as an alert, not as a complaint.