All case studies
NOC ServicesRepresentative engagement

Visibility Before Automation

A representative engagement, drawn from the environments ITBuilders works in across the Kingdom.

The context

Monitoring has become something organisations buy more than once. The first system is purchased to solve a visibility problem, produces alerts nobody can act on, and is quietly abandoned. The second is purchased to solve the first. What rarely changes between them is the underlying assumption — that monitoring is a tooling decision rather than a data one.

A logistics operator running distribution centres and a national fleet had reached the second system. Both were still installed. Neither was trusted. The operations team had learned that the dashboard showed green during outages and red during maintenance windows, so they had stopped looking at it and gone back to the method that worked: someone in a warehouse calls someone in IT.

The cost of that was invisible until it was measured. Incidents were reported by users an average of forty minutes after they began, and roughly a fifth were reported to the wrong team first. Mean time to restore was not a metric anyone could produce, because nobody agreed on when incidents had started.

The existing monitoring was not wrong in its readings. It was wrong in its scope. Availability was polled at device level across a network where almost nothing failed at device level. Links degraded. Circuits flapped. Applications slowed. None of that produces a down interface, so none of it produced an alert.

The approach

ITBuilders started with service definition. For each business function that mattered — warehouse scanning, dispatch, fleet telematics, back-office systems — the underlying dependency chain was documented, and monitoring was attached to the chain rather than its components. A scanning outage now registers as a scanning outage, not as four unrelated amber indicators across three consoles.

Thresholds were built from observed baselines over a defined measurement period rather than from defaults. Warehouse traffic patterns are extreme and predictable; what looks like saturation at 02:00 is normal at 02:00 and genuinely abnormal at 14:00. Alerting that cannot tell the difference will be ignored within a month, and correctly so.

Escalation paths were defined next, with each alert class carrying a named first responder and a documented action. Anything without a defined action was not enabled. That single rule removed more noise than any tuning exercise.

Only then was automation introduced, and narrowly — automated remediation for the small set of recurring faults with known, safe, reversible fixes. Circuit failover validation. Service restarts on a defined list. Configuration drift correction against approved templates. Everything else escalates to a human, because automation applied to a poorly understood environment produces faster mistakes rather than fewer ones.

What changed

The operator now detects the majority of incidents before users report them. Capacity decisions are made against trend data instead of the last complaint. The number that changed the internal conversation was not availability, which had always been claimed as high. It was the count of incidents identified by monitoring rather than by phone call, which moved from near zero to the clear majority within the first quarter of operation.

Continuity of operation

ITBuilders operates the NOC function from the Kingdom on a continuous basis — triage, threshold maintenance as traffic patterns shift, service definition for new business functions, and the trend reporting the operator's capacity planning now depends on. Monitoring that is built and handed over becomes the abandoned first system again within two years. The service definitions decay as the business changes, and maintaining them is the work.

Your next step

Facing a similar challenge?

Talk to ITBuilders about the constraints, priorities and operating requirements of your environment.

Start a conversation