Blog

One notification per fault, not one per ONU: root-cause alerting in FTTH networks

· 3 min read

When a PON port goes down, the NOC doesn't need 38 alerts. How to group symptoms under their cause and what information that single notification has to carry.

It's 3 a.m. and the on-call phone buzzes 38 times in a row. Each buzz is an ONU with no signal. They all hang off the same PON port. The technician doesn't need 38 notifications: they need one that says "PON port 1/2/3 on OLT North is down, 38 ONUs affected".

The problem with alerting on symptoms

Monitoring that alerts on every device that stops responding turns one fault into an avalanche. And that avalanche has concrete costs:

  • The technician takes longer to find the cause among so many notifications.
  • The important notifications get buried under the repeated ones.
  • Over time, the team starts muting notifications. And that's when the one that mattered gets missed.

The dependency chain

In a GPON network, each ONU depends on a chain of equipment. If one link goes down, everything hanging off it goes down too:

  1. The VPN tunnel to the ISP's MikroTik, if the OLT is reached through it.
  2. The OLT.
  3. The PON port or the NAP box.
  4. The subscriber's ONU.

Root-cause alerting means walking that chain from top to bottom and notifying only the link that failed, with a count of everything affected below it. If the tunnel is down, there's no point reporting that the OLT isn't responding: you already know why.

What the notification needs to say

A good root-cause alert answers, in one line, what the technician is going to ask:

  • What failed: the specific port, NAP box or OLT.
  • How much is affected: the number of ONUs behind it.
  • Why, if known: loss of signal (LOS) or loss of power (dying gasp).

That last distinction saves truck rolls. If most of the ONUs in a NAP box reported loss of power before going down, the most likely cause is a power outage in the area, not a fiber cut. That difference decides whether a field crew goes out or you wait for the power to come back.

Fast, but without false alarms

Speed matters: an outage detected in the 10-minute polling cycle is reported too late. If the OLT sends its alarms via syslog, an ONU going down is known instantly. But an ONU that's simply rebooting also reports that it went down, and it's not worth waking anyone up for that. A good rule is to wait one minute: if the alarm is still open, alert; if the ONU is back, don't.

Escalating without repeating

An alert that starts as a warning can become critical if the fault grows. That deserves one new notification, not one every five minutes. And when the equipment comes back, the alert closes on its own: nobody has to remember to mark it resolved.

And when it's maintenance

If the outage is scheduled, alerts should be logged but not notified, and that time shouldn't count against the SLA. A maintenance window set up in advance takes care of both.

How eCloud OLT Controller does it

The dashboard walks the chain VPN tunnel → OLT → port or NAP → ONU and notifies only the link that fails, via Telegram, email or signed webhook. It has more than 20 alert types, thresholds you can edit live and maintenance windows. What it doesn't do yet: it doesn't send SMS or WhatsApp, and the cause of each outage only shows up if the OLT reports it.

More detail on the platform page.

← All articles