Ekinops Dark Fiber Learning Path

Protection, Backup, Upgrade & Handoff — Operating for the Long Run

A service that turned up clean still has a life ahead of it: a fiber will be cut and protection must actually switch, a card will fail and must be swapped without losing config, software will need upgrading without bricking a live network, and eventually the whole thing must be handed to operations with proof it works. This page turns each of those into a runnable checklist — how to safely test a protection switch, how to back up and replace a card, how to upgrade with a rollback plan, how to verify licensing, and what evidence package proves the job is done.

Mental model Everything on this page is risk management around a live network. The recurring pattern is the same four moves: capture state before you touch anything, make one change at a time, verify against the captured baseline, and keep a way back. Protection testing, card replacement, and software upgrade are just that pattern applied to different components. Learn the pattern and each procedure becomes obvious.
Change-control discipline Nothing on this page is done outside an approved change window with a written rollback plan and impact analysis. Forcing a protection switch, pulling a card, or loading software on a production optical network are all service-risk events even when they "should" be seamless. No change without: an approved window, a captured pre-change baseline, a tested rollback, alarm suppression scoped to the expected events, and a named person watching for anything unexpected. "It's just a quick one" is how outages start.
Scope & accuracy note All CLI, menu names, alarm strings, licence identifiers, version strings, and readings here are representative/illustrative, in a generic format, and are not captured data. Exact protection schemes, switch commands, backup/restore workflow, upgrade procedure, and licensing model vary by platform and release — verify against current official Ekinops / Celestis documentation. Functional terms are used in place of model names; verify exact SKUs, capacities, and compatibility against current official Ekinops documentation.

Track your progress

Part 1 — Protection configuration & forced-switch testing

Protection only counts if it actually switches when the primary fails — and the only way to know is to test it deliberately, in a window, before a real fault does the testing for you. A protection scheme (optical-layer or OTN-layer — see the Protection & OTN page) has a working path and a protect path; a switch moves traffic from one to the other. Testing means forcing that switch on purpose and watching the hit.

How to safely trigger and verify a protection switch

Runnable checklist — forced protection switch test
Key idea — what to watch A protection test has three pass criteria: it switches (traffic actually moves to the protect path), it switches fast enough (within the scheme's spec — a hit, not an outage), and it reverts cleanly (no stuck state, no lingering alarm). The classic trap is switching onto a protect path you never verified — if the standby is quietly broken, your "test" becomes a real outage. Always confirm the standby is healthy before you force onto it.

Part 2 — Configuration backup, restore & card replacement

The iron rule: back up before you touch anything. A current configuration backup is what turns a failed card, a bad upgrade, or a fat-fingered change from a crisis into a restore. Backups are cheap; their absence is expensive exactly when you can least afford it.

Backup / restore

Runnable checklist — configuration backup

Replace a failed card — the sequence

Order matters. Do not pull a card before you know exactly what it was and can prove the replacement matches.

Runnable checklist — card replacement

Part 3 — Software upgrade with rollback

Upgrading software on a live optical network is the highest-stakes routine change you do. The whole discipline is proving compatibility first, upgrading in a staged order so one failure does not take everything, keeping a tested way back, and validating against baseline afterward — not just "it booted."

PhaseWhat you doGate to pass before proceeding
Compatibility checkConfirm the target release supports every installed card/optic and the NMS version; read release notes for known issues and required intermediate steps.Every element supported; upgrade path is valid.
BackupFull config backup + inventory of every node in scope, stored off-box.Backups verified readable.
Rollback planWritten, tested procedure to return to the current release/config, with the previous image staged and the criteria that trigger a rollback defined up front.Rollback documented and the previous image available.
Staged rolloutUpgrade a small, low-risk or lab/pilot node first; validate; then expand in controlled waves — never the whole network at once.Pilot passes full validation.
Post-upgrade validationVerify version, then verify SERVICES against baseline (power, FEC, alarms, PM), not just that the node is up.All services match baseline; no new alarms.
Runnable checklist — software upgrade

Part 4 — Licensing / entitlement verification

Some capacity, features, or ports may be gated by licences/entitlements. A change can silently fail or a feature can go dark if the licence does not cover it — so verify entitlement as part of any capacity change or upgrade, both before (will the change be allowed?) and after (did the upgrade preserve entitlements?).

Runnable checklist — licensing / entitlement

Part 5 — Acceptance, evidence package & operational handoff

A service is not "done" when the light is up — it is done when operations can run it without you, and you can prove it was delivered to spec. The acceptance/evidence package is that proof. Assemble it as you go; reconstructing it later is painful and always incomplete.

Document / evidenceWhat it proves
As-builtExactly what was installed and how it is wired: nodes, cards, serials, firmware, optics, wavelengths, patching — the reality, reconciled to the design.
Turn-up baselinesPer-service optical/PM readings at acceptance (power, OSNR, pre/post-FEC, ES/SES/UAS) — the reference for all future drift detection.
Test recordsLoss-budget/OTDR results for the span, service verification (soak) results, and protection-switch test times both directions.
RunbookHow to operate and recover the service: contacts, escalation, common alarms and their meaning, protection behaviour, backup/restore steps.
Inventory & sparesThe reconciled inventory and where spares for each card/optic type live.
Backups & licencesVerified config backups stored off-box and the entitlement record.
Change recordApprovals, windows, and the audit trail of who provisioned what.
Field checklist — acceptance & operational handoff