Ekinops Dark Fiber Learning Path

Operational Capstone — Provision, Prove, and Hand Off a Live Service

The design capstone on Link Design proved you can engineer a link on paper — close every budget, pick the gear, defend the reach. This capstone proves the other half of the job: that you can take a messy, half-true scenario packet and turn it into a running, tested, documented, defensible service. You will reconcile conflicting facts, build the BOM and channel plan, close the budgets for real in the Labs, provision it in Celestis, establish baselines, run an acceptance test, diagnose injected faults, prove protection with a forced switch, and hand off an as-built package a peer could operate from. Field readiness is not knowing the theory — it is producing the evidence.

Mental model Design answers "will it work?" Operations answers "is it working, how do you know, and can the next engineer keep it working after you leave?" A design that closes on paper but has no baseline, no acceptance record, and no runbook is not a delivered service — it is a liability with a green light on it. Everything below is built to force the evidence into existence.
This complements, does not repeat, the design capstone The full budget-closure method — the seven gates, the worked multi-span example, the scorecard — lives on Link Design, and the interactive budget tools live in the Labs. This page does not re-teach how to close a budget. It hands you a realistic packet and makes you use those pages to close it, then carries the closed design all the way to a signed-off, monitored, handed-off service. When a gate below says "close the budgets," you go compute them in the Link-Design Workbench and Engineering pages, then bring the numbers back here as evidence.

Capstone gates — track your progress

All artifacts below are representative / illustrative Every requirement, inventory quantity, OTDR trace, power reading, alarm, and PM figure on this page is invented for training and deliberately rounded or simplified. There are no real Ekinops SKUs — cards and optics are described by function and rate only. Real projects have vendor part numbers, real trace files, and real acceptance templates; use those in production and verify against current Ekinops / Celestis documentation. The point here is the method and the judgement, not the numbers.

The scenario packet — "Northgate Metro 100G, with a protect path"

This is your customer request, your site survey, your warehouse list, and your test data — exactly as they would arrive in the real world: from different people, on different days, not fully consistent. Read all of it before you touch a single gate. The planted contradictions are the assignment, not an accident.

Business & capacity requirements representative

  • Customer: Northgate Health, connecting a primary data centre (Site A) to a DR data centre (Site Z).
  • Service now: one 100GbE service, A↔Z, carrying storage replication (latency-sensitive).
  • Growth (18 mo): add a second 100GbE and two 10GbE management/backup services on the same fiber pair — plan the grid for this now.
  • Availability: 99.99% on the primary 100G; a fiber cut must not cause a hard outage — protection required.
  • Ownership: Northgate owns the routers/optics; you own everything from the client handoff inward.

Site & span info representative

  • A–Z working (short) path: stated on the customer order as 62 km, "single continuous span, no amps."
  • A–Z protect (long) path: quoted by the fiber provider as ~95 km over a diverse route, "fully diverse from the working path."
  • Fiber type: standard G.652 single-mode on both paths.
  • Patching: 2 connector pairs per path at each end (ODF + shelf), plus in-line splices per the OTDR.
  • Facilities: A and Z both have rack, power, and cooling; the protect path enters both buildings through the same entrance vault as the working path (noted by the site surveyor in passing).

Available Ekinops inventory on hand representative

This is what the warehouse says is physically on the shelf and allocatable to this job — function and rate only, no SKUs. You may not assume anything not on this list is available without a lead-time note.

Item (functional description)Qty on handNotes
Ekinops360 shelf + common/control + power (equipped)2One at A, one at Z; slots free.
Coherent 100G/200G-capable transponder card (1 client + 1 tunable DWDM line)2Enough for one A↔Z 100G line, both ends. None spare.
10G OTN muxponder card (10× 10G client → aggregated line)2For the two future 10G services; line side must land on the grid.
Passive DWDM mux/demux (fixed 8-channel, C-band)28 channels max per pair. One pair per site.
Tunable DWDM client optics (coherent, C-band)4Populate the transponder line ports.
Grey 100GbE client optics (short-reach, for the customer handoff)2Client side of the 100G transponder.
Optical protection switch module (1+1 line-side)1One only. Protects a single line direction pair.
Fixed-gain inline optical amplifier (EDFA, booster/pre-amp)0None on hand — lead time if required.
Variable optical attenuators (VOA, plug)4For receiver-overload / power-trim.

Representative OTDR trace readings representative

Fiber characterisation from the provider's acceptance sweep, 1550 nm, both paths, A→Z direction. Read the total end-to-end loss and the measured length — and compare the length against what the order form claimed.

PathOTDR measured lengthEventsTotal insertion loss (fiber+events, no connectors at ends)
Working (short)73.4 km2 splices (0.15 dB, 0.22 dB); 1 macrobend event (0.6 dB) at 41 km~17.1 dB
Protect (long)96.1 km4 splices (0.12–0.28 dB); no anomalies~20.4 dB

Representative per-connector power readings representative

Light-source/power-meter drop test across the mated end connectors that the OTDR sweep did not include (ODF and shelf bulkheads), plus one flagged reading.

ConnectorReadingExpectedFlag
A-end ODF bulkhead (working)0.3 dB<0.5 dBok
A-end shelf pigtail (working)0.4 dB<0.5 dBok
Z-end ODF bulkhead (working)1.7 dB<0.5 dBhigh — dirty?
Z-end shelf pigtail (working)0.4 dB<0.5 dBok

Existing alarm & PM history representative

The shelves were pre-staged and looped in the lab a week ago; the NMS carried that history over. It is not clean.

NMS event/PM extract (Site A shelf, last 7 days) [representative] [7d ago] INFO Card discovered: slot 3 coherent transponder, FW 4.2.1 [7d ago] MINOR slot 3 line: LOS (lab loopback removed) <- stale, from staging [6d ago] WARN slot 3 line: module temperature 71 C (soft high) <- recurred twice since [3d ago] MINOR slot 5 muxponder: client 1 LOS <- no client attached yet [2d ago] -- PM 15-min bins begin populating on slot 3 line [1d ago] TCA slot 3 line: pre-FEC crossed soft threshold, one 15-min bin [now] ACTIVE slot 3 line: module temperature 70 C (soft high) <- STILL PRESENT

Protection requirement representative

  • Primary 100G must survive a single fiber-path failure with automatic recovery (1+1 line protection is the intended mechanism).
  • Switch time and revertive behaviour to be defined and tested, not assumed.
  • Customer contract language says "fully diverse protect path."

The order-form fine print representative

  • "Working path is 62 km, no amplification required."
  • "Existing 8-channel mux is sufficient for all current and planned services."
  • "Protect path is fully diverse."
  • "Turn-up target: 5 business days from kit on-site."
The trap — at least four facts in this packet are wrong or misleading A field-ready engineer does not build what the order form says; they build what the evidence supports and document every place the two disagree. Before you open Gate 1, find the contradictions yourself. There are at least four. Hint categories: a distance that the OTDR contradicts (and what that does to the loss/OSNR budget and to the "no amp" claim); an inventory/grid item that cannot actually carry the full planned service set; a "diverse" protect path that is not truly diverse; and a power reading that will bite you at acceptance if you provision over it.

How the gates work

Ten gates, in order. Each is a gate with an objective, explicit acceptance criteria, a checklist you tick as you produce the evidence, and a collapsible "what good looks like" you open after you have attempted it. A gate is not done because you read it — it is done when the artifact it demands exists. Do them in sequence; later gates consume earlier gates' outputs.

Gate 1 — Reconcile the conflicting facts G1

Objective: Produce a one-page reconciliation register: every place the packet contradicts itself or the measured data, the resolution you chose, the evidence you trusted, and every assumption you are making explicit so a reviewer can challenge it.

Acceptance criteria:

  • Every planted contradiction is identified, with the two conflicting sources named.
  • Each is resolved by trusting measured evidence over stated claims, with a one-line rationale.
  • Assumptions that cannot yet be verified are listed as assumptions, not facts.
  • Any item that changes the design (distance → budget, diversity → protection value) is flagged for the relevant later gate.
ConflictSourcesResolution
Working-path lengthOrder form "62 km, no amp" vs OTDR 73.4 kmTrust the OTDR. Design to 73.4 km and ~17.1 dB fiber loss + connectors. The extra 11 km erodes loss and OSNR margin; re-run the budget before committing to "no amp."
Grid capacity"8-channel mux sufficient" vs service set = 2×100G + 2×10GTwo 100G lines take two channels; the muxponder aggregates the 10G clients onto one line channel. That is 3 channels — the 8-channel mux fits, but the claim that "8 channels is plenty" hides that each mux is a single point of failure per site and only one mux pair is on hand. Note the SPOF; it does not block turn-up but belongs in the risk register.
"Fully diverse" protect pathProvider quote vs surveyor's shared-entrance-vault noteNot truly diverse. The two paths share the building entrance vault at both ends — a single backhoe or vault flood takes out both. Protection still helps for mid-span cuts, but the contract word "fully diverse" is not met. Escalate: either accept documented shared-risk segments or get a truly separate entrance.
Z-end ODF connectorPower drop test 1.7 dB vs <0.5 dB specAlmost certainly a dirty or damaged endface. Inspect and clean (or replace) before turn-up. If you provision over it, it will look like span loss and poison your baseline.

Explicit assumptions (attackable): single fiber pair per path; connectors are the only un-swept loss; the coherent transponder's tunable optic reaches both budgets without an amp pending the Gate 3 recompute; the module-temperature warning is an environmental/staging artifact to be re-checked at baseline, not a card fault. The discipline: state what you assumed so the reviewer knows exactly where to push.

Gate 2 — BOM and channel plan G2

Objective: A build-of-materials that maps every on-hand item to a slot/port, and a channel plan that seats today's 100G plus the growth services on the C-band grid without collisions.

Acceptance criteria:

  • Every item is sourced from the on-hand inventory, or explicitly marked lead-time / to-order with the reason.
  • The channel plan assigns a distinct grid channel to each line signal and leaves room for the growth set.
  • The plan states which items are single-points-of-failure and which the protection module covers.
  • It is consistent with Gate 1 (e.g. it does not silently assume "no amp" if Gate 3 later needs one).

Reuse the Lab 2 channel-plan builder to place channels and prove no collisions. A workable plan (representative channel labels only):

ServiceCard @ A / ZClient opticLine opticGrid channelVia mux port
100G-#1 (today)Coherent transponder, slot 3Grey 100GbE SRTunable coherentCh AMux port 1
100G-#2 (growth)Coherent transponder (to-order — only 2 on hand, both used by #1)Grey 100GbE SRTunable coherentCh BMux port 2
10G ×2 (growth)10G OTN muxponder, slot 52× 10G greyAggregated line opticCh CMux port 3

To-order / gaps: a second pair of coherent transponders for 100G-#2 (inventory shows only 2, consumed by #1) — lead-time item, does not block today's turn-up. SPOFs: the single 8-channel mux/demux pair at each site (all services ride it); only one 1+1 protection module on hand, so only the primary 100G line is protected — the 10G aggregate and future 100G-#2 are unprotected unless more protection modules are ordered. Consistency with G1: if the Gate 3 recompute at 73.4 km fails OSNR without an amp, add a booster/pre-amp EDFA to the BOM as a to-order item now, so procurement is not surprised.

Gate 3 — Close the budgets (for real, in the Labs) G3

Objective: Close every gate the design capstone defines — loss, OSNR, dispersion, receiver power, capacity, and protection — for the as-measured working path (73.4 km) and the protect path (96.1 km), with numeric margins. Do not eyeball it; compute it.

Go compute, then bring numbers back Open the Lab 1 Link-Design Workbench and enter the measured distance, span count, and connector/splice counts from the packet; use the Link Engineering, Dispersion, Coherent Optics, and Capacity & Design pages for the per-budget method, and the Lab 3 protection-diversity designer for the protect path. The full ordered method and scorecard are on Link Design — this gate just makes you run it against the real, trap-laden numbers.

Acceptance criteria:

  • Loss budget closes on BOTH paths with the connector/splice reality (including the Z-end connector once remediated), with margin stated in dB.
  • OSNR closes at the coherent rate on both paths; if the 73.4 km (not 62 km) path fails without an amp, that is a finding, not a failure — document the amp requirement.
  • Dispersion within the coherent transponder's electronic-compensation tolerance at both lengths.
  • Receiver power lands inside the window (no overload at the short path, no starvation at the long path); VOA placement noted where needed.
  • Capacity: the channel plan fits the grid with growth headroom.
  • Protection: the protect path independently closes every budget, because after a switch it must carry the service alone.

The load-bearing insight: the order form's 62 km would have closed comfortably with margin to spare and justified "no amp." The measured 73.4 km plus a high Z-end connector eats roughly 3–4 dB more than planned. A typical outcome:

  • Working path (73.4 km): loss ≈ 17.1 dB fiber+events + ~1 dB clean connectors + mux insertion. Against a modern coherent Rx sensitivity this usually still closes on loss with a few dB margin — but only after the 1.7 dB Z connector is fixed. On OSNR at 100G coherent, a single unamplified 73 km G.652 span is generally fine; the finding is that you were one bad connector and 11 unplanned km away from needing a pre-amp, so record the true margin, don't inherit the order form's comfort.
  • Protect path (96.1 km, ~20.4 dB): this is the harder path. It may need a booster or pre-amp EDFA to hold OSNR margin at 100G — and there are zero EDFAs on hand (Gate 2). If the workbench shows the protect path failing OSNR without an amp, the protection design is not real until that amp is ordered and installed: a protect path that cannot carry the service after a switch is decoration.
  • Receiver power: short path may run hot into the Rx — place a VOA (4 on hand) to trim; long path must not starve — confirm it clears sensitivity with amp if required.
  • Dispersion: both lengths are well within a coherent transponder's electronic CD compensation range; no dispersion-compensating modules needed.

The lesson: the whole reach is set by the first gate that fails, and here the trap distance moved you from "trivially closes, no amp" to "closes on the working path but the protect path needs gear you don't have." That finding is the entire value of doing this gate against measured data instead of the order form.

Gate 4 — Define the service & OTN client-to-line mapping G4

Objective: Specify the service as the NMS will model it end to end, and define the OTN mapping for the 10G muxponder aggregate — which client lands in which ODU tributary on the line.

Acceptance criteria:

  • The 100G service is defined A-client → line channel → Z-client as one object, with rate and encoding stated.
  • The 10G services' OTN mapping is explicit: each 10G client → its ODU tributary → the aggregated line channel.
  • Protection is expressed as part of the service (which line pair is working, which is protect, switch mode).
  • The mapping is consistent with the Gate 2 channel plan and Gate 3 budgets.

See OTN Mapping and Protection & OTN for the mechanics. A clean service definition (representative):

Service "NGHEALTH-100G-1" (representative) A client port (grey 100GbE) --> transponder slot3 --> line Ch A | | [1+1 protection module] working line --> mux port1 --> WORKING path (73.4 km) | | +----------------------------------------------> protect line --> PROTECT path (96.1 km) | Z client port (grey 100GbE) <-- transponder slot3 <-- line Ch A Service "NGHEALTH-10G-AGG" (representative, growth) client 1 (10GbE) --> ODU2 trib #1 -. client 2 (10GbE) --> ODU2 trib #2 >-- muxponder slot5 --> line Ch C --> mux port3 -'

Protection expressed in the service: 1+1 line protection, working = short path, protect = long path, non-revertive chosen so a restored-but-not-yet-proven working fiber does not auto-carry traffic back before you have re-verified it (revertive vs non-revertive is a real decision — document why). The mapping matches Gate 2's Ch A/B/C assignment and Gate 3's per-path budgets, so nothing provisioned here contradicts what was proven to close.

Gate 5 — Provision it G5

Objective: Realise the service on the hardware through the NMS, following the standard provisioning workflow — not card-by-card CLI that the NMS cannot model.

Walk the provisioning steps Follow the ordered procedure on Operations · Provisioning (discovery → inventory reconcile → card/optic config → cross-connect → service creation → protection group → syncipate/verify). Provision through the service object so correlation and impact analysis work later; a hand-built cross-connect the NMS does not know about is invisible to every operational tool downstream.

Acceptance criteria:

  • Both shelves discovered; inventory reconciled against the Gate 2 BOM (right cards, right firmware, right optics in the right slots).
  • The 100G service is created as one NMS object, not a pile of manual cross-connects.
  • The 1+1 protection group is configured with working/protect paths and switch mode from Gate 4.
  • The service comes up green end to end and the pre-existing stale alarms (staging LOS, phantom muxponder client LOS) are cleared or explained.

Reconcile inventory first: the discovery must show the slot-3 coherent transponder at FW 4.2.1 (from the history) and the slot-5 muxponder — matching the BOM. Clear the staging artifacts deliberately: the 7-day-old line LOS was a removed lab loopback (clear it), the muxponder "client 1 LOS" is expected because no 10G client is attached yet (suppress/annotate, don't chase it). The module-temperature warning that is STILL ACTIVE is not a staging artifact — do not clear it blind; carry it into Gate 6/Gate 8 as a live item to explain. Create NGHEALTH-100G-1 as one service object, attach the protection group, and confirm the service view shows green A→Z on the working path with the protect path armed. If you built cross-connects by hand and the service view can't show the end-to-end object, you provisioned it wrong — redo it through the service layer.

Gate 6 — Establish operational baselines G6

Objective: Capture the golden-day readings that every future incident will be compared against, and set meaningful TCA thresholds off them. The day it turns up healthy is the most valuable day to record it.

Baseline against the monitoring page Use Operations · Monitoring for what to record and where, and Celestis NMS for the PM concepts (bins, ES/SES/UAS, pre/post-FEC, TCAs). Baselines feed Gate 7 acceptance and Gate 8 diagnosis — without them, "seems slow" can never become "we lost 3.5 dB."

Acceptance criteria:

  • Local and remote Tx/Rx power (dBm) recorded on both directions of the working path, and on the protect path.
  • Pre-FEC BER and confirmed zero post-FEC recorded; OSNR/Q where the platform exposes it.
  • Module temperatures recorded — including a decision on whether the active 70 C soft-high warning is normal-for-this-environment or a real problem.
  • TCA thresholds set relative to baseline (not vendor defaults), so they mean something.
  • Baseline stored in the NMS and in the acceptance record.

A representative baseline table (this is the artifact — a bare "-18 dBm" with no context is not a baseline):

PointWorking pathProtect path
A line Tx / Z line Rx+1.0 / -18.6 dBm+1.0 / -21.2 dBm
Z line Tx / A line Rx+1.0 / -18.9 dBm+1.0 / -21.4 dBm
Pre-FEC BER2e-96e-9
Post-FEC00
Slot-3 module temp70 C (soft-high present)

The temperature call: 70 C soft-high that recurs is a genuine finding — check rack airflow, filter, and slot adjacency before signing off; either remediate it or record an accepted-with-note deviation. Do not baseline a warned state as "normal" silently. TCAs: set pre-FEC and Rx-power thresholds a sensible margin off these baseline values (e.g. Rx-power TCA ~3 dB below baseline) so a real degrade trips a Warning before it becomes a customer hit — exactly the early-warning save shown on the Celestis page.

Gate 7 — Execute the acceptance test & record evidence G7

Objective: Prove the service meets spec with a repeatable acceptance test and capture the evidence a customer would sign against — not "link is up," but measured, thresholded, recorded.

Acceptance criteria:

  • An error-free transmission test (e.g. a defined-duration BER/traffic soak, or PRBS where the service can be held) with a pass threshold stated up front.
  • Optical readings compared against the Gate 6 baseline and within tolerance.
  • Latency measured on the working path and confirmed acceptable for the storage-replication use case.
  • Zero post-FEC, zero SES/UAS across the soak window; any ES explained.
  • Evidence captured as records (screenshots/exports/timestamps), not recollection.

See Turn-Up and Test & Measure for the procedures. A defensible acceptance record states the threshold before the result: e.g. "≥ 15-minute error-free soak, target BER better than 1e-12 post-FEC (i.e. zero uncorrected), Rx power within 1 dB of baseline, one-way latency < the replication budget." Then the recorded result: post-FEC 0, SES 0, UAS 0 over the window; Rx -18.6/-18.9 dBm matching baseline; latency consistent with 73.4 km of glass (~0.37 ms one-way as a sanity check, not the 62 km figure — another place the trap distance shows up). The critical discipline: a result with no pre-stated threshold is an anecdote; a threshold with no captured evidence is a claim. Acceptance needs both, filed.

Gate 8 — Diagnose the injected faults G8

Objective: Three faults are injected against the running service. For each, work it from the customer edge inward using presence-vs-quality and local-vs-remote comparison, name the root cause, and cite the reading that proves it.

Use the incident method Work each fault with the ladder from Troubleshooting and match the signature against the Incident Library. Do not jump to "it's the fiber" — let the readings localise it. The scenario snapshots below are representative.

Acceptance criteria: each fault has a named root cause, the one reading that distinguishes it from its look-alikes, and the correct action (fix vs hand-back-across-demarc vs schedule-a-window).

Fault A

Customer: "100G down." Service view: RED. Client Rx (A): present, 100GbE up. Local line Tx (A): +1.0 dBm. Local line Rx (A): DARK. Remote line Tx (Z): +1.0 dBm. Remote line Rx (Z): DARK. Span: LOS both dir, OSC down. Other waves on span: N/A (only wave).

Fault B

Customer: "100G flaky/slow." Service view: GREEN. Post-FEC: climbing. Pre-FEC: near correction limit. SES: accumulating. Local Rx (A): -22.4 dBm (baseline -18.9). Trend: -18.9 -> -22.4 over 30 h. Z-end ODF connector: recently reworked. Other waves: single wave.

Fault C

Customer: "no link at all." Service view: GREEN, optics nominal. Client Rx (A): DARK. Line Rx/Tx: all nominal, PM clean. Remote client Rx (Z): present. Customer changed their router optic this morning (per change log).

Fault A — bidirectional fiber cut on the working path. Distinguishing reading: both local Rx and remote Rx dark while both Tx are healthy, plus span LOS and OSC down. Light leaves both ends and dies in between, both directions → glass, not card. Action: this is exactly what protection is for — confirm/trigger the switch (Gate 9), then dispatch OTDR to the span. Root cause is the cut; the client-side and span alarms are its shadow.

Fault B — degrading connector / added loss, caught by PM not by an alarm. Distinguishing reading: service still GREEN with a steady one-direction Rx decline (-18.9 → -22.4 dBm) and pre-FEC eroding toward the limit with post-FEC just starting — the classic slope, not floor. The recently-reworked Z-end ODF connector (the 1.7 dB reading from the packet that you were told to fix) is the prime suspect: it was cleaned, but not verified. Action: schedule a window, inspect/clean/reseat that connector, re-verify against baseline — before post-FEC becomes a customer outage. This is the packet's connector trap coming back to bite exactly where predicted.

Fault C — customer-side fault beyond your demarc. Distinguishing reading: your line optics and PM are all nominal and the service object is green, but the client Rx at A is dark — there is no light arriving from the customer's router for the Ekinops to carry, and the change log shows they swapped their optic this morning. Action: hand back across the demarc with evidence (your client Rx dark, your line clean, baseline intact) — likely a wrong/failed customer optic or patch. Do not open a transport ticket; do not chase your own clean gear.

The through-line: presence vs quality vs direction, worked from the edge inward, tells all three apart. A is presence lost both directions in the glass; B is quality degrading on the trend while presence holds; C is presence lost outside your boundary. Same method, three very different actions.

Gate 9 — Protection forced-switch test G9

Objective: Prove the protection actually protects — force a switch, verify the service rides the protect path within spec, verify it carries traffic error-free there, and verify controlled return.

Acceptance criteria:

  • A forced/manual switch moves the service to the protect path; switch time measured and within the stated target.
  • On the protect path the service runs error-free (post-FEC 0) at the protect-path baseline optics from Gate 6 — proving Gate 3's "protect path independently closes" was real.
  • Return to working is controlled per the revertive/non-revertive decision from Gate 4, with re-verification before traffic rides the restored working fiber.
  • The shared-vault diversity finding from Gate 1 is restated in the test record — protection is proven against mid-span cuts, not against a shared-entrance event.

See Protection & OTN. A real forced-switch test, not a checkbox: trigger the switch from the NMS, confirm the service view shows traffic now on the protect (long) path, and measure the hit — a 1+1 optical switch should be sub-50 ms-class or per the module's spec; record the actual. Then hold it there and confirm post-FEC 0 and Rx at the protect-path baseline (-21.2/-21.4 dBm) — this is where a protect path that failed Gate 3's OSNR without an amp would betray you: it might switch fine but then run with errors, which is worse than an honest fail. Return per the non-revertive choice from Gate 4: switch back manually only after re-verifying the working fiber, so a repaired-but-flaky span doesn't silently reclaim traffic. Close by restating the honest limit: this protection defends a mid-span cut on either route, but because both routes share the entrance vault (Gate 1), it does not defend a vault event — the customer's "fully diverse" expectation is documented as not met.

Gate 10 — As-built package, runbook, and design defense G10

Objective: Produce the handoff a peer could operate from with no context transfer — the as-built, the runbook, and a short written defense of the decisions where you departed from the order form.

Lifecycle handoff Structure the package per Operations · Lifecycle so it plugs into ongoing operations, upgrades, and eventual decommission — the service you turn up today is something someone else inherits.

Acceptance criteria:

  • As-built reflects reality (measured 73.4 km, the actual slot/port/channel map, firmware) — not the order form's 62 km fiction.
  • Runbook covers baselines, TCA thresholds, the protection switch/return procedure, and the top-three fault signatures with their actions (from Gate 8).
  • Risk register carries the open items: shared-vault non-diversity, single-mux SPOF per site, protect-path amp requirement, to-order second transponder pair.
  • Design defense (half a page) explains each departure from the packet and the evidence behind it — the reconciliation register from Gate 1, made narrative.

The defense is the tell that separates a technician from an engineer. It reads like: "We built to the measured 73.4 km, not the ordered 62 km, because the OTDR is ground truth and the extra 11 km plus a bad Z connector cost ~3–4 dB; the working path still closes with recorded margin. We flagged that the protect path (96.1 km) needs a pre-amp we do not yet stock, so protection is armed but the amp is on order and the risk register carries the gap. We documented that the 'fully diverse' protect path shares both entrance vaults, so the contract term is not met — customer must accept the shared-risk segment or fund a separate entrance. We remediated the 1.7 dB Z connector before baseline; it re-degraded once (Fault B) and is the top watch item. Second 100G and the 10G pair fit the grid but need a to-order transponder pair and are unprotected with the single protection module on hand."

Every claim in that paragraph traces to a reading or a gate above. That is a defensible service — provisioned, proven, and honestly documented, warts included.


Final evidence package — the service is not "done" until every item exists

Evidence package checklist
Key idea The design capstone proved you can make a link close. This one proves you can make a service real and durable: reconciled against messy truth, provisioned so the tools can see it, baselined so the future has something to compare against, tested so the customer can sign, diagnosable when it breaks, protected in a way you have actually verified, and handed off so it outlives you. If you can produce every artifact in the checklist above from a packet that was lying to you in four places, you are field-ready.

Capstone complete?

All ten gates ticked at the top of this page, and every item in the evidence-package checklist produced. Then take it back to the Labs to stress the growth case, and to Self-Check to confirm the concepts.