Troubleshooting — Isolate by Blast Radius
Optical troubleshooting is fast once you stop guessing and start narrowing. The single most powerful question is: how many things are affected? One client? One wavelength? Every wavelength on a fiber? A whole shelf? The size of the blast radius tells you which layer to look at and lets you skip the rest. This page gives you the ladder as a reasoning tool, five problem buckets each with its next diagnostic action, the pre-FEC/post-FEC and directional-fault reasoning that separates a slow degrade from an outage, and the data set to collect before you escalate.
Track your progress
Reading the method is not the same as running it. The Incident Library puts you through 16 graded, evidence-driven incidents — no-light, one-way, dirty connector, rising pre-FEC, one wavelength vs all wavelengths, protection-switch failure, timing faults and more. It scores you on scoping blast radius first, gathering the discriminating evidence before you touch hardware, and reaching the right root cause efficiently. Work the ladder below, then go prove it there.
The blast-radius ladder
Walk from the top (smallest, most specific) down. Wherever the scope of impact matches, that is your layer. Everything below that rung is shared infrastructure; everything above is customer-specific. The ladder is a reasoning tool: the count of affected things IS the diagnosis of the layer, before you read a single power level.
The ladder as a reasoning tool — worked scope examples
Read the count, then predict the layer. Practise until it is reflex:
| What the customer/NOC reports | Scope you confirm | Layer it points to | First confirming check |
|---|---|---|---|
| "My one 10G circuit is down"; every other service on the same span is clean | ONE client / ONE wave | Client side, or that wave's service/card | Is the client handoff dark but transport green? → client. Transport also alarming for that wave only? → transponder/config. |
| Three customers on the same fiber all down within the same second | ONE span (every wave) | Physical fiber / line system | LOS + OSC down across the span, other spans clean → the glass. |
| Every service on node C down, neighbours reachable | ONE shelf | Hardware / power / controller | Is the shelf reachable in the NMS at all? Power/controller/common-card alarm? |
| Everything at the Raleigh site dark at once | ONE site | Facility / fiber entrance | Site power? Fiber entrance / OSP into the building? |
| Multiple sites degrade together, no clean cut | WHOLE system | Correlated line-system / mgmt / protection | Amplifier chain, protection state, management-plane event across the correlated path. |
Five problem buckets
Once the ladder points you at a layer, match the symptoms to one of these buckets. Each lists what you see, the usual causes, and — most importantly — the next diagnostic action and the reading that confirms or denies it. Do not skip to a fix; take the confirming reading first.
1. Physical fiber issue
Symptoms: loss of light / loss of signal; a large sudden Rx drop vs baseline; a slow multi-day Rx decline (creeping loss); an OTDR event at a specific distance; OSC down; multiple wavelengths on the same span impacted together.
Causes: fiber cut, bad/unseated patch cord, dirty connector endface, wrong strand patched, bad/high-loss splice, macro-bend (pinched cable, over-tight coil, tray crush), damaged/contaminated ferrule.
Next action → reading: Read local Rx vs baseline and compare to remote Tx. Rx far below baseline while remote Tx is normal → loss is in the span (that direction). Then OTDR/OLTS the span → an event/high-loss point at a distance confirms the physical fault and locates it.
Tell: the blast radius is the whole span. If every wave on the fiber degrades or dies together, it is the glass.
2. Client-side issue
Symptoms: Ekinops line side healthy and far-end optics fine, but the client signal is missing; router/switch interface down; "my port won't come up"; client CRC/encoding errors.
Causes: customer router/switch interface down or errored, wrong/failed optic on the client device, speed/encoding mismatch (expecting 100GbE, getting something else), dirty/bad client patch, wrong client optic type.
Next action → reading: Read client-side Rx on the Ekinops port. No client Rx → the customer is not transmitting into you; the fault is theirs (their optic/patch/interface) — hand back across the demarc with the reading. Client Rx present but no signal recovered → check rate/encoding match.
Tell: the transport is green end to end; only the handoff is dark. The problem lives on the customer's side of the demarc.
3. Wavelength / service issue
Symptoms: exactly one channel impacted while every other wavelength on the same fiber runs clean; that one wave shows rising pre-FEC or intermittent hits.
Causes: wrong channel/wavelength assigned or tuned, failed/degrading transponder or muxponder, laser drift on a single tunable optic, aging optic (creeping laser bias), service misconfiguration for that wave.
Next action → reading: Confirm neighbours on the same glass are clean (rules out fiber), then read that card's local Tx and laser bias/temperature. Local Tx low/absent → transmitter fault (RMA the card). Tx fine but far end sees nothing on that wave only → wrong channel/tuning; verify the assigned channel.
Tell: neighbours on the same glass are healthy, so it is not the fiber. One wave misbehaving points at the card or the service for that wave.
4. Line-system issue
Symptoms: multiple wavelengths degraded (not necessarily dead), amplifier / ROADM / OSC / power alarms, span tilt or total-power alarms, a group of services on the same line all showing worse pre-FEC together.
Causes: EDFA amplifier fault or automatic laser shutoff, ROADM/WSS switch fault, OSC down, gain/tilt misadjustment, upstream power problem cascading down the line.
Next action → reading: Read the amplifier states and per-channel power/tilt along the span. An amp in automatic shutoff → trace UPSTREAM (its shutoff is usually a safety response to a break above it). Gain/tilt out of spec with a group of waves degraded together → line-system adjustment, not a single card.
Tell: a group of waves degrade together but it is not a clean cut — the shared active line gear is the common element.
5. Configuration / NMS issue
Symptoms: light is clearly present and healthy (good Rx power, clean FEC) but the service is not passing traffic; or the service came up wrong right after a change.
Causes: incorrect cross-connect, wrong OTN mapping (ODU into the wrong tributary), protection in the wrong state (locked out / forced to a failed path), mismatched service parameters between the two ends.
Next action → reading: Confirm good light + clean FEC (proves physics), then read the cross-connect / OTN mapping / protection state and diff it against the design and the audit log. A recent audit entry at the failure time → someone changed it; reconcile and revert.
Tell: physics is fine, logic is wrong. Good light + no service = look at the config and the NMS, not the fiber.
Pre-FEC vs post-FEC — read the margin, act before the outage
This is the interpretation skill that separates engineers who prevent outages from those who only clean them up. Modern optics run FEC: the line is expected to arrive with errors, and FEC corrects them. So a non-zero pre-FEC BER is normal and healthy. What matters is the margin — how far pre-FEC sits below the FEC correction limit — and its direction of travel.
| What PM shows | Customer impact | What it means | Action |
|---|---|---|---|
| Pre-FEC steady, well below limit; zero post-FEC | None | Healthy, full margin | Baseline it; nothing to do |
| Pre-FEC rising over hours/days; still below limit; zero post-FEC | None yet | Margin eroding — something is degrading (loss creeping, OSNR falling) | Act in a window. Inspect/clean, reseat, investigate the span BEFORE it hits the limit |
| Pre-FEC at/over the limit; post-FEC > 0; SES climbing | Real hits now | Out of margin — FEC can no longer correct; customer is affected | Incident. Work it now; expect an SLA event |
| Post-FEC bursty/intermittent | Intermittent hits | Marginal path — dirty connector, marginal OSNR, thermal drift | Find the marginal element; do not wait for a hard failure |
Directional-fault reasoning — local Rx vs remote Tx
Fiber faults are directional. A pair carries two independent directions; a cut, dirty connector, or bad splice can hit one direction and leave the other clean. If you only read your local side you will misread a one-direction fault as "everything." The technique: compare each receiver against the transmitter that feeds it, at both ends.
How to use it: take the affected receiver's reading, then jump to the far end and read the transmitter that should be feeding it. If the transmitter is healthy and the receiver is dark/low, the loss is between them, in that direction — which tells the fiber crew which strand/path to chase and rules out "both my cards are bad." If both receivers are dark while both transmitters are healthy, you have a bidirectional event (a real cut or a mid-span fault). Always confirm the far-end Tx before concluding "no light = cut."
Dispatch or fix it remotely?
- Fix remotely (no truck) when the evidence points at logic or a soft state you can change from the NMS: wrong cross-connect / OTN mapping, protection stuck in the wrong state, a config that changed in the last window (audit log), or a far-end element you can reset/re-provision. Good light + bad service = almost always remote.
- Dispatch when the evidence is physical and directional: Rx far below baseline with a healthy far-end Tx, an OTDR event at a distance, OSC down across a span, a dirty/needs-reseat connector, or a card that must be reseated/RMA'd. Bad light = hands on glass.
- Localise before you roll. Use directional reasoning to send the crew to the RIGHT end/segment — dispatching to A when the loss is in the Z→A direction on a mid-span splice wastes a truck roll. The OTDR distance and the "which direction" answer tell you where.
- If protected, confirm whether traffic already switched to the standby path (customer may be fine) — that changes dispatch from emergency to scheduled, and lets you work the failed path in a proper window.
Worked decision flow
- Open the service view. Is light present at the local line Rx? No → check the far end's Tx. Far end transmitting but you see nothing? Physical fiber (bucket 1) or line system (bucket 4) — check whether neighbours are also affected to decide.
- Light present but service not passing? Config/NMS (bucket 5). Check cross-connect, OTN mapping, protection state, and the audit log.
- Light present and quality bad (high pre-FEC, climbing post-FEC/errored seconds)? If only this wave: wavelength/service (bucket 3). If several waves on the span: physical (bucket 1, e.g. dirty connector) or line system (bucket 4).
- Transport green end to end but the customer handoff is dark? Client side (bucket 2) — their optic/interface.
Report: NOC pages: three customers on the RAL↔RTP fiber all hard down, same minute. Alarm view shows ~40 Critical/Major alarms.
Reason it through before revealing: what is the blast radius, what is the smallest shared element, and what one reading confirms it?
Report: One customer's 100G service takes intermittent hits. Every other wavelength on the same fiber is clean. PM shows pre-FEC on that wave has climbed over two days; post-FEC just went from zero to occasional bursts. Local Rx has drifted 3 dB below baseline; remote Tx is normal.
Reason it through: which bucket, which direction, dispatch or not?
Report: A service that worked yesterday is not passing traffic this morning. Line Rx is healthy and within baseline, pre-FEC is clean, zero post-FEC. There was a maintenance window overnight on that node.
Reason it through: physics vs logic — where do you look?
Before you call Ekinops support, collect this data
A support case moves at the speed of the data you bring to it. Collect all of this first so the first response is analysis, not "please send us the following." Each line notes why support needs it — understanding the why makes you collect it correctly.
Try it: troubleshooting simulator
A 100G wave is down. Work it like a NOC engineer — scope the blast radius, localize by direction, and read the data before swapping hardware. Each choice gives feedback and the path ends in a diagnosis.