Ekinops Dark Fiber Learning Path

Troubleshooting — Isolate by Blast Radius

Optical troubleshooting is fast once you stop guessing and start narrowing. The single most powerful question is: how many things are affected? One client? One wavelength? Every wavelength on a fiber? A whole shelf? The size of the blast radius tells you which layer to look at and lets you skip the rest. This page gives you the ladder as a reasoning tool, five problem buckets each with its next diagnostic action, the pre-FEC/post-FEC and directional-fault reasoning that separates a slow degrade from an outage, and the data set to collect before you escalate.

Mental model: This is the same discipline as triaging a data-centre outage. "One VM is down" and "the whole cluster is down" send you to completely different places. In optics the blast radius maps directly onto physical layers: the more services affected, the lower and more shared the failed component. Let the scope of impact pick your layer — then confirm with one targeted reading before you touch anything.

Track your progress

Product note: Alarm names, PM screens, and menu paths differ by Celestis/Ekinops release — verify exact behaviour and labels against current Ekinops / Celestis documentation. The reasoning here (blast radius, presence vs quality, directionality, FEC margin) is platform-independent and is what actually resolves incidents.
Practice this

Reading the method is not the same as running it. The Incident Library puts you through 16 graded, evidence-driven incidents — no-light, one-way, dirty connector, rising pre-FEC, one wavelength vs all wavelengths, protection-switch failure, timing faults and more. It scores you on scoping blast radius first, gathering the discriminating evidence before you touch hardware, and reaching the right root cause efficiently. Work the ladder below, then go prove it there.

The blast-radius ladder

Walk from the top (smallest, most specific) down. Wherever the scope of impact matches, that is your layer. Everything below that rung is shared infrastructure; everything above is customer-specific. The ladder is a reasoning tool: the count of affected things IS the diagnosis of the layer, before you read a single power level.

SCOPE OF IMPACT MOST LIKELY LAYER WHERE TO LOOK ───────────────────────────────────────────────────────────────────────────────── ONE CLIENT / handoff only → client side / customer → client optic, patch, router IF │ (their side of demarc) ▼ ONE WAVELENGTH (others on wavelength / service / transponder/muxponder, same fiber are FINE) → config → cross-connect, OTN mapping │ ▼ ONE FIBER SPAN (EVERY physical fiber / the glass: cut, dirty/bad wavelength on it impacted) → line system → connector, splice, bend, strand │ ▼ ONE SHELF (all services on hardware / power / shelf power, controller, that node down) → controller → common card, chassis │ ▼ ONE SITE (every shelf at a facility / power / site power, fiber entrance, location down) → fiber entrance → building, OSP into the site │ ▼ WHOLE SYSTEM (multiple sites network-wide / mgmt / protection state, mgmt plane, / everything) → correlated event → correlated line-system failure
Key idea: "Is it just this one wavelength, or every wavelength on the fiber?" is the highest-value single question in optical troubleshooting. One wave down = go UP the ladder to service/config. Every wave down = go DOWN the ladder to the physical span. Answer that before you touch anything.

The ladder as a reasoning tool — worked scope examples

Read the count, then predict the layer. Practise until it is reflex:

What the customer/NOC reportsScope you confirmLayer it points toFirst confirming check
"My one 10G circuit is down"; every other service on the same span is cleanONE client / ONE waveClient side, or that wave's service/cardIs the client handoff dark but transport green? → client. Transport also alarming for that wave only? → transponder/config.
Three customers on the same fiber all down within the same secondONE span (every wave)Physical fiber / line systemLOS + OSC down across the span, other spans clean → the glass.
Every service on node C down, neighbours reachableONE shelfHardware / power / controllerIs the shelf reachable in the NMS at all? Power/controller/common-card alarm?
Everything at the Raleigh site dark at onceONE siteFacility / fiber entranceSite power? Fiber entrance / OSP into the building?
Multiple sites degrade together, no clean cutWHOLE systemCorrelated line-system / mgmt / protectionAmplifier chain, protection state, management-plane event across the correlated path.
Reasoning shortcut: blast radius and "shared-ness" move together. The more services fall together, the lower and more shared the failed component. One customer = their edge. One wave = that wave's card/config. One span = the shared glass. One shelf = shared hardware. One site = shared facility. You are really asking "what is the smallest thing all the victims have in common?" — that common element is the fault.

Five problem buckets

Once the ladder points you at a layer, match the symptoms to one of these buckets. Each lists what you see, the usual causes, and — most importantly — the next diagnostic action and the reading that confirms or denies it. Do not skip to a fix; take the confirming reading first.

1. Physical fiber issue

Symptoms: loss of light / loss of signal; a large sudden Rx drop vs baseline; a slow multi-day Rx decline (creeping loss); an OTDR event at a specific distance; OSC down; multiple wavelengths on the same span impacted together.

Causes: fiber cut, bad/unseated patch cord, dirty connector endface, wrong strand patched, bad/high-loss splice, macro-bend (pinched cable, over-tight coil, tray crush), damaged/contaminated ferrule.

Next action → reading: Read local Rx vs baseline and compare to remote Tx. Rx far below baseline while remote Tx is normal → loss is in the span (that direction). Then OTDR/OLTS the span → an event/high-loss point at a distance confirms the physical fault and locates it.

Tell: the blast radius is the whole span. If every wave on the fiber degrades or dies together, it is the glass.

2. Client-side issue

Symptoms: Ekinops line side healthy and far-end optics fine, but the client signal is missing; router/switch interface down; "my port won't come up"; client CRC/encoding errors.

Causes: customer router/switch interface down or errored, wrong/failed optic on the client device, speed/encoding mismatch (expecting 100GbE, getting something else), dirty/bad client patch, wrong client optic type.

Next action → reading: Read client-side Rx on the Ekinops port. No client Rx → the customer is not transmitting into you; the fault is theirs (their optic/patch/interface) — hand back across the demarc with the reading. Client Rx present but no signal recovered → check rate/encoding match.

Tell: the transport is green end to end; only the handoff is dark. The problem lives on the customer's side of the demarc.

3. Wavelength / service issue

Symptoms: exactly one channel impacted while every other wavelength on the same fiber runs clean; that one wave shows rising pre-FEC or intermittent hits.

Causes: wrong channel/wavelength assigned or tuned, failed/degrading transponder or muxponder, laser drift on a single tunable optic, aging optic (creeping laser bias), service misconfiguration for that wave.

Next action → reading: Confirm neighbours on the same glass are clean (rules out fiber), then read that card's local Tx and laser bias/temperature. Local Tx low/absent → transmitter fault (RMA the card). Tx fine but far end sees nothing on that wave only → wrong channel/tuning; verify the assigned channel.

Tell: neighbours on the same glass are healthy, so it is not the fiber. One wave misbehaving points at the card or the service for that wave.

4. Line-system issue

Symptoms: multiple wavelengths degraded (not necessarily dead), amplifier / ROADM / OSC / power alarms, span tilt or total-power alarms, a group of services on the same line all showing worse pre-FEC together.

Causes: EDFA amplifier fault or automatic laser shutoff, ROADM/WSS switch fault, OSC down, gain/tilt misadjustment, upstream power problem cascading down the line.

Next action → reading: Read the amplifier states and per-channel power/tilt along the span. An amp in automatic shutoff → trace UPSTREAM (its shutoff is usually a safety response to a break above it). Gain/tilt out of spec with a group of waves degraded together → line-system adjustment, not a single card.

Tell: a group of waves degrade together but it is not a clean cut — the shared active line gear is the common element.

5. Configuration / NMS issue

Symptoms: light is clearly present and healthy (good Rx power, clean FEC) but the service is not passing traffic; or the service came up wrong right after a change.

Causes: incorrect cross-connect, wrong OTN mapping (ODU into the wrong tributary), protection in the wrong state (locked out / forced to a failed path), mismatched service parameters between the two ends.

Next action → reading: Confirm good light + clean FEC (proves physics), then read the cross-connect / OTN mapping / protection state and diff it against the design and the audit log. A recent audit entry at the failure time → someone changed it; reconcile and revert.

Tell: physics is fine, logic is wrong. Good light + no service = look at the config and the NMS, not the fiber.

Common mistake: Assuming "loss of light" always means a fiber cut. Loss of light also comes from a far-end laser that shut down, an amplifier in automatic shutoff, or a Tx/Rx swap in the patch path. Confirm the far-end is transmitting before you dispatch a fiber crew — the NMS shows you remote Tx in seconds.

Pre-FEC vs post-FEC — read the margin, act before the outage

This is the interpretation skill that separates engineers who prevent outages from those who only clean them up. Modern optics run FEC: the line is expected to arrive with errors, and FEC corrects them. So a non-zero pre-FEC BER is normal and healthy. What matters is the margin — how far pre-FEC sits below the FEC correction limit — and its direction of travel.

What PM showsCustomer impactWhat it meansAction
Pre-FEC steady, well below limit; zero post-FECNoneHealthy, full marginBaseline it; nothing to do
Pre-FEC rising over hours/days; still below limit; zero post-FECNone yetMargin eroding — something is degrading (loss creeping, OSNR falling)Act in a window. Inspect/clean, reseat, investigate the span BEFORE it hits the limit
Pre-FEC at/over the limit; post-FEC > 0; SES climbingReal hits nowOut of margin — FEC can no longer correct; customer is affectedIncident. Work it now; expect an SLA event
Post-FEC bursty/intermittentIntermittent hitsMarginal path — dirty connector, marginal OSNR, thermal driftFind the marginal element; do not wait for a hard failure
Key idea: Pre-FEC is your early-warning slope; post-FEC is the alarm floor. Rising pre-FEC with zero post-FEC is a gift — the customer feels nothing while you still have a maintenance-window's worth of time to fix the cause. The moment post-FEC goes non-zero, that window has closed and you are in an outage. Watch the trend, not just the current value.

Directional-fault reasoning — local Rx vs remote Tx

Fiber faults are directional. A pair carries two independent directions; a cut, dirty connector, or bad splice can hit one direction and leave the other clean. If you only read your local side you will misread a one-direction fault as "everything." The technique: compare each receiver against the transmitter that feeds it, at both ends.

A-SITE SPAN Z-SITE ───────── ───────── A.Tx ───────────────────────────────────────────────────► Z.Rx (direction A→Z) A.Rx ◄─────────────────────────────────────────────────── Z.Tx (direction Z→A) READ ALL FOUR: Z.Rx low but A.Tx normal → loss in the A→Z direction (that strand/path) A.Rx low but Z.Tx normal → loss in the Z→A direction (the OTHER strand/path) BOTH Rx dark, both Tx normal → bidirectional cut / mid-span line-system fault A local seems up, Z.Rx dark, A.Tx normal → suspect a Tx/Rx swap at A

How to use it: take the affected receiver's reading, then jump to the far end and read the transmitter that should be feeding it. If the transmitter is healthy and the receiver is dark/low, the loss is between them, in that direction — which tells the fiber crew which strand/path to chase and rules out "both my cards are bad." If both receivers are dark while both transmitters are healthy, you have a bidirectional event (a real cut or a mid-span fault). Always confirm the far-end Tx before concluding "no light = cut."

Dispatch or fix it remotely?

Decision — dispatch a truck vs fix from the NMS:
  • Fix remotely (no truck) when the evidence points at logic or a soft state you can change from the NMS: wrong cross-connect / OTN mapping, protection stuck in the wrong state, a config that changed in the last window (audit log), or a far-end element you can reset/re-provision. Good light + bad service = almost always remote.
  • Dispatch when the evidence is physical and directional: Rx far below baseline with a healthy far-end Tx, an OTDR event at a distance, OSC down across a span, a dirty/needs-reseat connector, or a card that must be reseated/RMA'd. Bad light = hands on glass.
  • Localise before you roll. Use directional reasoning to send the crew to the RIGHT end/segment — dispatching to A when the loss is in the Z→A direction on a mid-span splice wastes a truck roll. The OTDR distance and the "which direction" answer tell you where.
  • If protected, confirm whether traffic already switched to the standby path (customer may be fine) — that changes dispatch from emergency to scheduled, and lets you work the failed path in a proper window.

Worked decision flow

  1. Open the service view. Is light present at the local line Rx? No → check the far end's Tx. Far end transmitting but you see nothing? Physical fiber (bucket 1) or line system (bucket 4) — check whether neighbours are also affected to decide.
  2. Light present but service not passing? Config/NMS (bucket 5). Check cross-connect, OTN mapping, protection state, and the audit log.
  3. Light present and quality bad (high pre-FEC, climbing post-FEC/errored seconds)? If only this wave: wavelength/service (bucket 3). If several waves on the span: physical (bucket 1, e.g. dirty connector) or line system (bucket 4).
  4. Transport green end to end but the customer handoff is dark? Client side (bucket 2) — their optic/interface.

Report: NOC pages: three customers on the RAL↔RTP fiber all hard down, same minute. Alarm view shows ~40 Critical/Major alarms.

Reason it through before revealing: what is the blast radius, what is the smallest shared element, and what one reading confirms it?

Scope: every wave on one span → physical fiber / line system (down the ladder). Shared element: the RAL↔RTP span. Confirming reads: local Rx dark at both ends of the span while the far-end Tx reads normal, OSC down across the span, neighbouring spans clean. That is a span event, and both-Rx-dark points to a bidirectional cut. Action: OTDR the span to locate the event distance, then dispatch to the located point (not blindly to a site). The 40 alarms are one root cause — the client-side link-downs at all three customers are leaves that clear when the fiber is restored. One incident, not forty.

Report: One customer's 100G service takes intermittent hits. Every other wavelength on the same fiber is clean. PM shows pre-FEC on that wave has climbed over two days; post-FEC just went from zero to occasional bursts. Local Rx has drifted 3 dB below baseline; remote Tx is normal.

Reason it through: which bucket, which direction, dispatch or not?

Bucket: neighbours clean rules out the fiber as a whole → wavelength/service (3) or a localised physical issue on that path. Direction: local Rx 3 dB low with remote Tx normal = loss in the far→local direction on that path. Pre/post-FEC: rising pre-FEC with post-FEC now bursting = margin just ran out; you are entering an incident (the early window is closing). Action: this is physical and directional (creeping loss on one direction) — likely a dirty/loosening connector or a degrading splice on that path. Dispatch to inspect/clean/reseat the connectors in that direction, or OTDR to locate; do not just RMA the card (the card's Tx is fine — the loss is in the glass). Had you caught it a day earlier at rising-pre-FEC/zero-post-FEC, it would have been a scheduled clean, not a hit.

Report: A service that worked yesterday is not passing traffic this morning. Line Rx is healthy and within baseline, pre-FEC is clean, zero post-FEC. There was a maintenance window overnight on that node.

Reason it through: physics vs logic — where do you look?

Bucket: good light + clean FEC + no traffic = Config/NMS (5). Physics is fine; logic is wrong. Confirming read: the audit log shows a change in last night's window; diff the current cross-connect / OTN mapping / protection state against the design. Likely a cross-connect pointed at the wrong tributary, or protection left locked out on a failed path. Action: fix it from the NMS — no truck. This is the highest-yield reconciliation in the packet: "recent changes" versus the failure time. Revert or correct the mapping and the service returns.

Before you call Ekinops support, collect this data

A support case moves at the speed of the data you bring to it. Collect all of this first so the first response is analysis, not "please send us the following." Each line notes why support needs it — understanding the why makes you collect it correctly.

Field checklist — support data packet (with the why):
Key idea: "Recent changes" is the highest-yield line in that packet. Most "sudden" outages trace to work done in the last maintenance window — a re-patch, a firmware push, a cleaning that disturbed a neighbor. Always reconcile the failure time against your change log before assuming a random hardware fault.

Try it: troubleshooting simulator

A 100G wave is down. Work it like a NOC engineer — scope the blast radius, localize by direction, and read the data before swapping hardware. Each choice gives feedback and the path ends in a diagnosis.