Design Labs — Build the Whole Path
Reading is not the same as designing. These six labs build on each other: a two-site 100G link, then more services, more wavelengths, a third site, protection, and finally operating it all through the NMS. Do them in order — each reuses the artifacts you made in the last. Work each one on paper (or a whiteboard) before you check the answers.
Lab completion
Each lab below is a live workbench: change the inputs, watch the page compute and grade your design against explicit acceptance criteria, and read the hints when a gate fails. A lab only ticks complete when its objective is genuinely met. The original written scenarios, worked budgets and reference artifacts are preserved under each lab’s “Worked solution” toggle. Labs 4 and 5 stay design exercises — they lead with an objective and reuse the Lab 2 and Lab 3 tools to prove the design.
Lab 1 — Two-site 100G link: design workbench
Deliver a 100G service over 80 km with ≥3 dB margin on BOTH the loss and OSNR budgets and dispersion within tolerance. Start from the failing preset and tune the workbench until every gate passes.
- Optic set to 100G coherent and distance = 80 km
- Loss budget closes with ≥3 dB margin
- OSNR closes with ≥3 dB margin
- Accumulated dispersion within tolerance
Lab 1 — Two-site 100G dark fiber design
Scenario: Site A and Site B are 40 km apart, connected by a single-mode dark fiber pair you lease. Each site has a router with a 100G Ethernet port. You must carry 100GbE between the two routers over Ekinops transport.
Requirements: one 100GbE service, A to B, over the 40 km pair, using Ekinops360 equipment. Unprotected is acceptable for this lab. Assume standard G.652 single-mode fiber, LC/APC connectors at the panels, and a coherent 100G line optic on a ~24 dB system budget.
Questions to answer:
Think through the flow, then reveal.
Client side = the 100GbE handoff from the customer router into the Ekinops (grey/short-reach optic). Line side = the DWDM-tuned wavelength from the Ekinops transponder/muxponder out onto the 40 km fiber. Same structure mirrored at both A and B: client Rx/Tx faces the router, line Rx/Tx faces the fiber.
Span loss ≈ fiber attenuation (order of ~0.25 dB/km at 1550 nm → ~10 dB for 40 km) plus connector and splice losses at each junction, plus any patch-panel hops. Sum the span loss and compare against the line optic's receiver sensitivity, keeping a margin (commonly a few dB) for aging and repair splices. If the budget closes with margin, 40 km is a passive span (likely no amplifier needed); if it does not, you add gain or pick a longer-reach optic. Verify exact attenuation and sensitivity numbers against the actual fiber and optic datasheets. The full worked numbers are in the reference solution below.
At 40 km for a single 100G wave, typically no — the span is short enough to be passive if the budget closes, and coherent 100G optics tolerate dispersion electronically. Confirm against the optic's reach spec; do not assume.
Do the lab on paper first, then open each artifact below and compare. These are the deliverables a design review would expect — not just answers, but the diagram, the numbers, and the runbook.
Key: λ1 = one DWDM wavelength on the ITU C-band grid. The pair is two strands — one Tx, one Rx — so the service is bidirectional. At 40 km the optical line system (OLS) is passive: mux/demux only, no amplifier. Tx at A must land on Rx at B and vice-versa; a crossed pair is the classic "power is present but link won't come up" fault.
| Element | Basis | Loss (dB) |
|---|---|---|
| Fiber attenuation | 40 km × 0.25 dB/km @ 1550 nm | 10.0 |
| Connectors | 4 × 0.3 dB (panel mates end-to-end) | 1.2 |
| Splices | 8 × 0.1 dB (fusion splices along route) | 0.8 |
| Mux + demux | passive WDM filter insertion loss | 3.5 |
| Design margin | aging, repair splices, dirt | 3.0 |
| Total required | sum of the above | 18.5 |
| System budget | line optic capability | 24.0 |
| Remaining margin | 24.0 − 18.5 | 5.5 |
Verdict: PASS. The path needs 18.5 dB (including 3 dB of built-in margin) against a 24 dB system, leaving 5.5 dB of headroom on top. That comfortably closes, so 40 km stays a passive span — no amplifier. Run the same table for B→A; on a symmetric pair the number is the same. Every figure here is a teaching approximation — verify attenuation, connector/splice loss, and the optic's real budget against the actual datasheets.
| Site | Shelf / slot | Port | Side | Optic / signal | Connects to |
|---|---|---|---|---|---|
| A | Muxponder, slot 1 | C1 | Client | 100GbE grey (short reach) | Router A, 100G port |
| A | Muxponder, slot 1 | L1 | Line | DWDM λ1 (coherent) | Mux/demux A, ch λ1 |
| A | Mux/demux A | COM | Line | Combined C-band | Tx/Rx fiber pair → B |
| B | Mux/demux B | COM | Line | Combined C-band | Tx/Rx fiber pair → A |
| B | Muxponder, slot 1 | L1 | Line | DWDM λ1 (coherent) | Mux/demux B, ch λ1 |
| B | Muxponder, slot 1 | C1 | Client | 100GbE grey (short reach) | Router B, 100G port |
Slot/port labels are illustrative — read the real ones off the actual chassis. The point is that every physical hop has a named endpoint at both ends, so a tech can patch it and a NOC can reference it on a ticket. λ1 must be the same ITU channel provisioned at both A and B.
- Fiber acceptance — OTDR/OLTS the pair both ways; confirm measured loss ≤ ~15.5 dB (18.5 budget minus the 3 dB margin) before touching gear.
- Inspect & clean — scope every endface; both ends LC/APC; never mate APC↔UPC.
- Install & cable — seat muxponders; patch client (C1) to router, line (L1) to mux/demux, COM to the outside pair. Verify Tx→Rx orientation.
- Provision — set λ1 (same ITU channel both ends), line rate/format, and the service object in the NMS.
- Verify power — read line Rx at both ends; confirm it sits inside the optic's expected Rx window and matches the budget prediction.
- Verify quality — confirm pre-FEC is low/zero and there are no post-FEC errors; check both directions.
- Confirm client — router link up, traffic passing error-free.
- Baseline & document — record local/remote Tx/Rx (dBm), pre-FEC, λ1, slots/ports, and file the acceptance record.
Work from the fiber up: power present but out of window = span/optic loss; power in window but errors = quality; all optical-clean = the fault is above the transport layer.
Expected design artifacts: a physical diagram (routers, shelves, client/line ports, the fiber pair); a loss budget worksheet; a client/line port map; a turn-up checklist; a troubleshooting decision tree. The five reference blocks above are exactly these five artifacts.
- Loss budget closes with margin — ~18.5 dB required vs a 24 dB system, ≈5.5 dB spare on top of a 3 dB design margin.
- Line Rx sits inside the optic's expected window and matches the budget prediction at both ends.
- Pre-FEC is low and there are zero post-FEC errors.
- Margin held is > 3 dB so the link survives aging and a future repair splice.
- Both directions (A→B and B→A) are healthy, not just one — check the far-end Rx too.
- Diagram shows correct Tx→Rx orientation; the port map names every slot/port; each troubleshooting rung maps to a layer.
Lab 2 — Channel plan: fit the services onto the grid
Place 3×100G and 1×400G onto the fixed DWDM grid with no channel collisions and the 400G on a valid wide (≈75 GHz, two-slot) allocation. Click a service, then click a starting slot.
- All four services placed on the grid
- No channel collision (no two services share a slot)
- The 400G sits on a valid wide (two-slot) allocation
Lab 3 — Add a second 100G wavelength
Scenario: Demand grows; you need a second, independent 100G service across the same A–B fiber pair.
Requirements: a second 100GbE service on a different wavelength over the existing pair, without disturbing the first. If Lab 1 already used a mux/demux for λ1, this lab is about adding λ2 to the same filters and re-checking the plan.
A mux/demux at each end to combine the two wavelengths onto the one fiber (and split them at the far end). Assign the second service a distinct DWDM channel on the ITU grid that does not collide with the first. The fiber now carries two colors; the mux is what lets them share the strand.
You add the insertion loss of the mux and demux to every wavelength's budget. Re-run the budget with the added passive loss; confirm both channels still close with margin. If they no longer do, that is where amplification enters the design.
Pick λ2 from the same DWDM plan the mux/demux supports (ITU-T G.694.1 C-band, e.g. 100 GHz spacing), on a channel the filter actually passes to a distinct port — you can't just tune to any frequency, it must match a filter port. A collision means two services on the same channel (they'd interfere and neither works); adjacent channels are fine as long as each optic stays on its grid point and doesn't drift. Record λ1 and λ2 in the channel-plan register so the next person doesn't reuse a taken channel.
Expected design artifacts: updated physical diagram with mux/demux carrying two colors; a channel plan table (service → ITU channel/frequency → mux port) for λ1 and λ2; revised loss budgets for both channels including mux/demux loss.
- No channel collision — λ1 and λ2 are distinct ITU channels on the supported grid.
- Both budgets close with the mux/demux loss included; if not, amplification is called out explicitly.
- Adding λ2 did not disturb the in-service λ1 (a make-before-break plan on the mux).
- The channel plan is written down in the register so future adds don't reuse a channel.
Related reference — grooming lower-rate clients (old Lab 2)
Lab 2 — Add two 10G services onto the 100G transport path
Scenario: The A–B path from Lab 1 exists. Two new customers each need a 10GbE service between A and B. You do not want to light new wavelengths for them.
Requirements: carry 2× 10GbE plus room for the existing traffic, reusing the current fiber/wavelength where possible. Assume the existing 100G is a real 100GbE client already filling the wave — read Q3 before you assume there's room.
Use a muxponder / OTN grooming: multiplex the 10G clients into ODUs and carry them inside the 100G line container (an ODU4 can carry multiple lower-rate ODUs). The two 10G clients become client ports on a muxponder; their traffic rides the existing wavelength as separate ODU2 tributaries. No new fiber, no new wavelength.
You add two client ports (10G optics) at each end, and two new service objects in the NMS, each mapped to a tributary in the OTN structure of the existing line. The line side does not add ports; the multiplexing is logical.
The optical layer is unchanged — same λ1, same 40 km span, same ~18.5 dB budget — because you're adding logical tributaries inside the existing wave, not new light. The real constraint is container capacity: an OTU4/ODU4 carries ~100G of client. If a full 100GbE already fills it, there is no room for 2× 10G and you need a muxponder that grooms (e.g. 10×10G into the 100G) rather than a straight 100G transponder — or you go to Lab 3 and light a second wavelength. Always confirm the tributary-slot map before promising the capacity.
Expected design artifacts: updated port map with the two new 10G client ports; an OTN mapping sketch showing the 10G ODU2 tributaries and their tributary slots inside the ODU4; two new service records; a note confirming the container has free tributary slots.
- You can articulate ODU-into-ODU multiplexing (ODU2 tributaries inside an ODU4).
- You verified there are free tributary slots before committing — not just assumed room.
- The two new 10G services are independently monitored (own PM, own service ID).
- The existing 100G service is untouched; the optical budget and wavelength are unchanged.
Lab 3 — Protection & diversity designer
Build a protected A→C 100G service: choose a working path and a physically diverse protect path (sharing no conduit) whose budgets both close (loss ≤ 21 dB on the 24 dB system).
- Working path budget closes (≥3 dB margin)
- Protect path budget closes (≥3 dB margin)
- The two paths are physically diverse (no shared conduit)
Goal: a protected A→C 100G service on a 24 dB system where both the working and protect paths close (≥3 dB margin → loss ≤ 21 dB) and are physically diverse.
The diversity trap. Two paths can look diverse on the logical map yet ride the same physical conduit. Here edge A–B and edge D–C both ride duct-1. So the tempting pair — working A→B→C (duct-1, duct-2) and protect A→D→C (duct-3, duct-1) — is not diverse: a single backhoe on duct-1 severs both at once. True diversity means the two paths share no conduit, not merely no logical edge.
A correct answer. Working A→B→C (duct-1, duct-2, 12 dB) with protect A→C direct short (duct-4, 12 dB): no shared conduit, and both close with 12 dB of margin. (Swapping which is working vs protect is equally valid.)
The budget trap. The A→C long diverse route (duct-6) is physically diverse from everything — but at 24 dB on a 24 dB system it leaves 0 dB margin, so it fails the ≥3 dB gate. A protect path is useless if it cannot actually carry the service. Check both gates: diversity and budget.
- Both budgets close — the protect path (often longer) is proven on its own, not assumed.
- No shared conduit — separate ducts/handholes/entrances, verified against physical records, not two circuit IDs.
- After a switch the service runs on its only remaining path — the NMS must alarm "running unprotected" so the cut is repaired urgently.
Lab 4 — Add a third site with add/drop
Insert Site C as an add/drop node: drop one wavelength to a local client, express the rest, and prove the A–B express budget still closes once the node’s insertion loss is added.
- Every wavelength has a clear disposition at C — labeled add/drop or express, nothing ambiguous.
- The dropped service has real client ports at C (shelf, slot, optic).
- The express A–B budget is re-run with the node loss and still closes with margin.
- You stated the OADM-vs-ROADM trade-off (loss/cost vs software flexibility) and justified the choice.
Reuse the Lab 2 channel-plan builder to keep C’s drop wavelength off the express channels, and re-run the Lab 1 workbench with an extra mux/demux pair (or higher connector/splice count) standing in for the node’s express insertion loss to confirm the pass-through budget still closes.
Lab 4 — Add a third site with add/drop
Scenario: A new Site C sits along the A–B route. Some services must drop at C; others must pass straight through to the far end.
Requirements: C drops one service locally and expresses the rest; existing A–B services keep working. C sits partway along the route, so the A–C and C–B spans are each shorter than the original A–B, but the node itself adds loss.
An OADM (fixed) or ROADM (reconfigurable) at C. It drops the chosen wavelength(s) to local client cards and expresses the remaining wavelengths through on the line. A ROADM lets you choose add/drop-vs-express per wavelength in software instead of re-patching.
The service destined for C is add/drop (terminates on a local client). The A–B services are express/pass-through (they transit C untouched on the line). Draw each wavelength and label it add/drop or express at C.
It adds the OADM/ROADM express insertion loss to every pass-through wavelength, on top of the (now split) fiber loss A–C plus C–B. Even though the total fiber distance is unchanged, the added node loss eats into your margin — re-run the express budget and confirm it still closes. If the added node pushes A–B under margin, that node is where an amplifier earns its place. Fixed OADMs generally add less loss than a full ROADM, which is a real trade-off vs the ROADM's software flexibility.
Expected design artifacts: a three-site (A–C–B) diagram with the OADM/ROADM at C; a per-wavelength table labeling add/drop vs express at each site; updated loss budgets for the A–C, C–B, and end-to-end A–B express paths including the node loss.
- Every wavelength has a clear disposition at C — labeled add/drop or express, nothing ambiguous.
- Express services are unaffected and their re-run budget still closes with the added node loss.
- The dropped service has real client ports at C (a shelf, slot, and optic — not just a line on the diagram).
- You stated the OADM vs ROADM trade-off (loss/cost vs software flexibility) and why you chose one.
Lab 5 — Add ring protection
Protect the critical service so a single span cut causes at most a sub-second hit (≈50 ms). Prove physical diversity of the two ring directions and that the longer protect path’s own budget closes.
- A documented single-cut survival story — name the cut, name the surviving direction.
- Verified physical diversity — separate ducts/handholes/entrances, not just two circuit IDs.
- The protect path’s own budget closes despite being longer than the working path.
- A plan to test and time the switch (≈50 ms) and an alarm for running unprotected afterwards.
Prove the two ring directions share no conduit with the Lab 3 protection-diversity designer, then close the longer protect path in the Lab 1 workbench by entering its greater distance and span count.
Lab 5 — Add ring protection
Scenario: Sites A, B/C, and a new diverse path form a ring. A key service must survive any single span cut.
Requirements: protect the critical service so a single fiber cut causes at most a sub-second hit (target ~50 ms); verify path diversity. The protect path is longer than the working path, so it has its own budget to prove.
Ring (or 1+1 over the two ring directions). The service is bridged both ways around the ring; the receiver selects the healthy direction. On a span cut, traffic wraps/steers the other way and the receiver switches with minimal loss.
True physical diversity: the two directions must not share a conduit, handhole, or building entrance, or a single backhoe kills both. Also confirm the protect path's budget closes on its own — it is useless if it cannot actually carry the service.
The switch itself should complete on the order of ~50 ms — fast enough that most upper-layer sessions ride through. But once switched, the service is on its only remaining path with no protection left: the NMS must raise an alarm for "protection unavailable / running on protect path" so someone treats the still-broken working span as urgent. Protection that switches silently and no one repairs the cut is a single fault away from a hard outage. Decide revertive (auto-return when the working path heals) vs non-revertive up front.
Expected design artifacts: a ring diagram showing working and protect directions; a diversity statement (routes, ducts, building entrances); the protect-path loss budget; a protection-switch test plan (how you'll force and time the switch).
- A documented single-cut survival story — name the cut, name the surviving direction.
- Verified physical diversity — separate ducts/handholes/entrances, not just two circuit IDs.
- The protect path's own budget closes despite being longer than the working path.
- A plan to test and time the switch (≈50 ms target) and an alarm for running unprotected after it fires.
Lab 6 — Operate & troubleshoot: graded incident
Diagnose a live incident from the service view: scope the blast radius and localize the fault before touching hardware, and reach the correct diagnosis in ≤3 steps for the top grade.
- Reached the correct diagnosis (fiber cut on the shared span)
- Scoped the blast radius before touching hardware
- Localized the fault without swapping/reseating hardware first
Lab 6 — Operate the service through Celestis NMS
Scenario: The multi-site, protected network is live. Now operate it: provision, baseline, monitor, and run an incident.
Requirements: use the NMS to provision one service end to end, capture its baseline, and walk an outage using service correlation.
Baseline optics (local/remote Tx/Rx in dBm), baseline PM (pre/post-FEC, errored seconds), the service ID, wavelength/channel, and module/slot/port at both ends — stored in the NMS and the acceptance record. Without the baseline, later degradation is un-diagnosable.
The NMS maps the service to its ports, wavelength, and spans, so one fiber alarm immediately tells you which customer services are affected and how big the blast radius is. You troubleshoot the service view, not dozens of raw card alarms — walk the "100G wave is down" steps from the Celestis page.
Run impact analysis on the node/card/span you plan to touch: the NMS lists every service that rides it, so you know the real blast radius before you pull anything, and can tell which customers to notify and which services are currently unprotected (Lab 5). Compare live optics against the captured baseline first, so you can prove the service was healthy going in — and detect if your change moved a power level or pre-FEC.
Expected design artifacts: a provisioned service object; a saved baseline (optics + PM); an impact-analysis output for one node; a written incident walk-through using the blast-radius ladder.
- You can provision a service end to end in the NMS and see it as one object, not scattered ports.
- A baseline is saved (Tx/Rx dBm, pre-/post-FEC) and referenced when reading live state.
- You run impact analysis before a change and can name the blast radius.
- You drive an incident from the service view — closing the loop from design (Lab 1) to operations.
Further practice
The design-review rubric and three stand-alone worked case studies below are unchanged reference material. Hold every lab deliverable against the rubric, then try each case study on paper before opening its worked solution.
Design review rubric
Hold every lab deliverable (and every real design) against these four gates before you call it done. If any one fails, the design is not ready — it will bite you at turn-up or on the first fault.
Every wavelength's budget summed and compared to the system budget, with a real margin (> ~3 dB). Amplification is called out where a span or added node doesn't close. Both directions checked.
Every service on a distinct ITU channel the filters actually pass — no collisions. Channels recorded in the register. Add/drop vs express disposition is unambiguous at every node.
Protected services ride physically diverse paths (separate ducts/entrances), the protect path's own budget closes, and the NMS alarms when a service is left running unprotected after a switch.
Turn-up records local/remote Tx/Rx (dBm), pre-/post-FEC, service ID, wavelength, and slot/port — stored in the NMS and the acceptance record so later degradation is diagnosable.
Case study A — Metro 100G ring with add/drop + protection
Scenario: Four sites — A, B, C, D — sit on a fibre ring, each span ~25 km (ring circumference ~100 km). You must deliver one 100G service from A to C. Node B is an add/drop node that also terminates its own local service; the A→C wave expresses through B. The A→C service must survive any single span cut (ring / SNCP protection).
Requirements: a per-hop and end-to-end loss budget for the working path, the worst-case OSNR around the ring, a channel plan that does not collide with B's local service, a protection path with defined switch behaviour, and a "what good looks like" acceptance bar. Standard numbers: 0.25 dB/km @1550, connectors 0.3 dB, splices 0.1 dB, mux/demux ~3.5 dB (terminal pair), design margin ~3 dB, ~24 dB system budget, doubling the number of spans costs ≈3 dB OSNR.
1 · Working-path loss budget (A→B→C, ~50 km, λ1 express through B):
| Element | Basis | Loss (dB) |
|---|---|---|
| Fiber attenuation | 50 km × 0.25 dB/km | 12.5 |
| Terminal mux/demux | add at A + drop at C | 3.5 |
| Express node at B | OADM/ROADM pass-through (illustrative ~1.5) | 1.5 |
| Connectors | 6 × 0.3 dB (panels A, B in/out, C, ends) | 1.8 |
| Splices | 6 × 0.1 dB along the route | 0.6 |
| Design margin | aging, repair splices, dirt | 3.0 |
| Total required | sum of the above | 23.9 |
| System budget | coherent 100G line optic | 24.0 |
| Remaining margin | 24.0 − 23.9 | 0.1 |
Verdict: PASS, but only just. The express node at B is the margin-eater — pass-through node loss stacks on top of the 3 dB design margin and leaves almost nothing. That is the lesson: on a ring, every express node you transit erodes the budget even though the fibre distance is unchanged. If you cannot live with ~0.1 dB spare, you either add a small amplifier at B, use a lower-loss fixed OADM instead of a full ROADM, or accept that this ring wants amplification. Re-run the identical table for the protect path A→D→C — same 50 km, same one express node (D) — so it closes to the same ~0.1 dB and must be proven independently.
2 · Worst-case OSNR around the ring: the two paths are each 2 spans. Relative to a single 25 km span, going from 1 span to 2 spans is one doubling → about −3 dB OSNR on either the working or protect path. The worst case is simply "the service is on a 2-span path," ~3 dB down from a single-span reference. The saving grace: the A→C service is 100G DP-QPSK, which needs the lowest OSNR of any coherent format, so a short metro ring has comfortable OSNR headroom even though the loss budget is tight. (Loss margin and OSNR margin are different gates — check both.)
3 · Channel plan (no collision):
| Service | ITU channel | Disposition at each node |
|---|---|---|
| A→C 100G (protected) | λ1 | add A · express B · express D · drop C — bridged both ring directions |
| B local service | λ2 (distinct) | add/drop at B only |
λ1 ≠ λ2 — distinct ITU channels the ring filters actually pass, recorded in the channel-plan register. Because SNCP bridges λ1 both ways around the ring, λ1 must be free on every span of the ring, not just the working path.
4 · Protection path + switch behaviour: this is ring / SNCP (1+1 over the two ring directions). The A→C signal is bridged onto both the A→B→C and A→D→C directions; C's receiver continuously selects the healthy one. On a cut of (say) span B–C, the working direction fails and C's tail-end selects the protect direction A→D→C. Target switch time on the order of ~50 ms. After the switch the service is on its only remaining path — running unprotected — so the NMS must raise a "protection unavailable" alarm so the cut span is repaired urgently. Decide revertive vs non-revertive up front.
- Both the working (A→B→C) and protect (A→D→C) budgets are computed and each closes — with the express-node loss explicitly included, not forgotten.
- You flagged the express node as the margin-eater and named the fix (amp, lower-loss OADM, or accept amplification) when spare margin is thin.
- No channel collision: λ1 (A→C) and λ2 (B local) are distinct, and λ1 is free on every ring span because it is bridged both ways.
- Protection is real: two physically diverse ring directions, ~50 ms switch target, and a "running unprotected" alarm after a switch.
- You checked both gates — loss budget (tight) and OSNR (comfortable on 100G QPSK) — and did not confuse the two.
Case study B — Long-haul 400G DCI over amplified spans
Scenario: Two data centres ~600 km apart, connected over an amplified line of ~8 × 75 km EDFA spans. You must carry 400G between them. A colleague asks "why can't we just use a cheap direct-detect 400G optic?"
Requirements: show quantitatively why direct-detect fails (dispersion and OSNR), estimate the accumulated OSNR after 8 spans, choose a coherent format (16QAM vs QPSK reach trade-off), and note whether a 400ZR+ pluggable or a full transponder is the right tool. Standard numbers: 0.25 dB/km, CD ~17 ps/nm·km, doubling the spans costs ≈3 dB OSNR.
1 · Why direct-detect fails — dispersion: accumulated chromatic dispersion over the whole line is
A direct-detect receiver has no DSP to reverse dispersion — it only sees brightness. At 400G line rates the pulses are so short that even a small fraction of 10,200 ps/nm smears each pulse across many neighbours; the signal is unrecoverable long before 600 km (direct-detect 10G already struggles past ~60–80 km unaided, and 400G is far more sensitive). Without dispersion-compensating modules (which coherent lines deliberately omit) direct-detect 400G is a non-starter here.
2 · Why direct-detect fails — OSNR: across 8 EDFA spans, noise accumulates. Using the doubling rule, 8 spans = 2³, i.e. three doublings from a single span → about −9 dB OSNR versus one span. A worked single-span estimate (illustrative, launch 0 dBm, EDFA NF ~5 dB, span loss 18.75 dB):
A direct-detect 400G scheme has none of coherent's DSP/FEC gain, so a ~25 dB end-of-line OSNR combined with 10,200 ps/nm of uncompensated dispersion leaves it no chance. Coherent is required.
3 · Coherent format choice — 16QAM vs QPSK: the accumulated OSNR is ~25 dB (illustrative).
| 400G option | OSNR need | Reach behaviour over 8×75 km |
|---|---|---|
| DP-16QAM (~64 Gbaud) | Higher | Spectrally efficient (fits ~75 GHz) but OSNR-hungry — ~25 dB is marginal; this format is happier at DCI distances (~120 km class), not 600 km. |
| DP-QPSK (higher baud / 2 carriers) | Lower | Needs the least OSNR → reaches 600 km comfortably, but at half the bits/symbol, so 400G-QPSK needs more baud or two carriers and more spectrum. |
So over 600 km you lean toward a QPSK-based 400G (or PCS-tuned near-QPSK) with strong SD-FEC (~11 dB+ coding gain), trading spectral efficiency for the reach the OSNR budget can actually support. 16QAM would need the line shortened, fewer spans, or a mid-line regen.
4 · 400ZR+ vs transponder: plain 400ZR (fixed 16QAM, ~120 km amplified) cannot do 600 km. A ZR+ pluggable — selectable modulation, lower rate for reach, stronger FEC — may reach it and collapses the box into the router (IPoDWDM), but you then own the OSNR/reach engineering across the amplified line. A full transponder gives the broadest format/reach range and keeps a clean IP/optical demarc. For a 600 km amplified DCI, verify the specific ZR+ reach spec against this OSNR estimate; if it is marginal, the transponder is the safer tool.
- You quantified both direct-detect failure modes — 10,200 ps/nm of uncompensated CD and ~9 dB of OSNR erosion over 8 spans — not just hand-waved "it's too far."
- The OSNR accumulation is estimated (~25 dB end-of-line) and tied to the doubling rule, so the number is defensible.
- The format choice is justified by the OSNR budget: QPSK/strong-SD-FEC for 600 km reach, with 16QAM correctly flagged as marginal at this distance.
- You made an explicit ZR+ vs transponder call and said you would verify the pluggable's real reach spec against the estimate before committing.
Case study C — SAN extension for DR (Fibre Channel over WDM)
Scenario: Two data centres run storage replication for disaster recovery over a WDM link carrying Fibre Channel. Management wants to know whether replication should be synchronous (every write waits for the far site to acknowledge) or asynchronous (writes acknowledge locally, replicate in the background) at two candidate distances: 40 km and 300 km.
Requirements: work the latency at both distances and make the sync-vs-async call; address buffer-credit / distance considerations for Fibre Channel throughput; and specify protection / diversity for the DR link. Standard number: light in fibre adds ~4.9 µs/km one-way.
1 · Latency at each distance (one-way = distance × 4.9 µs/km; a synchronous write costs a full round trip):
| Distance | One-way | Round trip (per sync write) |
|---|---|---|
| 40 km | 40 × 4.9 = 196 µs | ≈ 392 µs (~0.39 ms) |
| 300 km | 300 × 4.9 = 1470 µs | ≈ 2940 µs (~2.94 ms) |
Decision: at 40 km, each synchronous write adds only ~0.39 ms of round-trip latency — usually acceptable, so synchronous replication is viable (zero data loss, RPO = 0). At 300 km, every write would stall ~2.94 ms waiting for the far-end ACK; on a write-heavy workload that throttles application throughput badly, so you move to asynchronous replication (writes ACK locally, replicate in the background) and accept a small non-zero RPO. The distance sets the mode: sync near, async far. (The exact cut-over depends on the application's write-latency tolerance — ~40 km sync / ~300 km async is the teaching split, not a hard law.)
2 · Buffer-credit / distance considerations: Fibre Channel uses buffer-to-buffer (BB) credits to keep frames flowing before an acknowledgement returns. The number of frames "in flight" grows with distance × line rate, so a link that is fine at 40 km can collapse in throughput at 300 km if the FC ports (or the WDM transport's distance-extension / credit-spoofing feature) do not provide enough BB credits to fill the longer pipe. At 300 km there are ~7.5× more frames in flight than at 40 km, so BB-credit sizing (or a transport that spoofs credits over distance) is essential — otherwise the link idles waiting for ACKs regardless of how much raw bandwidth the wavelength has.
3 · Protection / diversity: a DR link is worthless if a single backhoe severs it at the same time you need it. Route the replication wavelength over a physically diverse path (separate conduits, handholes, building entrances) and consider a protected wavelength so a single cut does not take DR down. For async replication a brief protection switch (~50 ms) is harmless — the queue drains afterward. For sync replication, a switch or any latency change directly stalls production writes, so diversity and a fast, well-tested switch matter even more, and you must confirm the protect path's own latency does not push sync past the application's tolerance.
- The latency math is shown at both distances (round trip ~0.39 ms @40 km, ~2.94 ms @300 km) and drives the call — sync at 40 km, async at 300 km.
- You tied the mode to RPO: sync = zero data loss but latency-bound; async = small RPO but distance-tolerant.
- Buffer credits are addressed — you noted throughput collapses over distance without enough BB credits / transport credit-spoofing, independent of raw bandwidth.
- The DR link is physically diverse and protected, and you noted the protect path's latency must not break a sync design.
Simulators
Interactive, in-browser practice tools embedded across the learning path. Each link jumps to the page that hosts that simulator — sweep the inputs and watch the verdicts move.
Span count, loss, NF and launch power vs required OSNR — live margin and PASS/FAIL. Open on Link engineering →
Accumulated CD vs rate tolerance, with coherent auto-pass. Open on Dispersion →
Draw and interpret clean / dirty-connector / bend / cut traces. Open on Fiber testing & loss budget →
Branching "100G wave down" fault tree that rewards localizing by direction. Open on Troubleshooting →
Loss + OSNR + dispersion gates with a scorecard and limiting gate. Open on Link design →