Catch Advisors
Network

SD-WAN Renewal: Retest Failover After Every Circuit Change

Your SD-WAN design passed failover testing when it was installed. Good.

Since then, a carrier replaced a circuit. Bandwidth changed. A branch moved to broadband. Firewall policies were updated. Voice moved to a new platform. Someone adjusted the path thresholds after users complained about performance.

The old test does not prove the current network.

SD-WAN renewal should include a controlled failover test after material circuit and policy changes. If the provider cannot show how today’s applications move across today’s paths, you are renewing a diagram and a dashboard. You are not renewing proven continuity.

Start with the changes since the last clean test

Do not begin with the renewal quote. Begin with the change record.

Build a list of everything that could have changed traffic detection, path selection, capacity, security inspection, or recovery:

  • Carrier, access type, bandwidth, handoff, or public IP changes
  • New circuits, disconnected circuits, or temporary links that became permanent
  • SD-WAN edge, firewall, modem, or cellular hardware replacements
  • Firmware and software upgrades
  • Routing, quality-of-service, application-classification, or security-policy changes
  • New cloud regions, SaaS platforms, voice services, VPNs, or allowlists
  • New monitoring targets, thresholds, timers, and alert routes
  • Changes to the managed service provider, carrier escalation path, or site contacts

Then ask a blunt question: which of these changes were followed by an end-to-end failover test?

A change ticket marked complete proves that the change was implemented. It does not prove that a voice call, payment transaction, VPN session, warehouse workflow, or cloud application survived the new path.

NIST’s contingency-planning guidance treats testing and maintenance as part of keeping recovery capability usable as systems and organizations change. That principle fits SD-WAN renewal. The evidence should move with the network.

Separate physical diversity from policy failover

SD-WAN can steer traffic across multiple links. It cannot make two links physically independent.

Before testing policy, confirm what changed underneath it. A new carrier name may still use the same last-mile owner, conduit, building entrance, upstream facility, power source, or local equipment as the primary circuit.

Use the internet circuit diversity buying guide to verify the transport layer. Record the carrier, last-mile owner, service ID, demarcation, public IPs, bandwidth, access type, building path, power dependency, and support contact for every link.

Keep two decisions separate:

  1. Can one physical event interrupt both paths?
  2. If one usable path degrades or fails, will the SD-WAN policy move the right traffic to the remaining path?

Passing the second test does not fix a shared trench. Buying diverse circuits does not fix a weak steering policy. You need evidence for both.

Test degradation, not only a dead cable

Unplugging the primary handoff is useful. It proves one failure condition. Real network trouble is not always that clean.

A circuit may stay electrically up while packet loss rises, latency jumps, jitter damages voice, an upstream route fails, or a cloud security path stops responding. The edge still sees a live interface. Users see a broken application.

Current vendor documentation shows why the details matter. Fortinet’s SD-WAN performance SLA documentation, for example, describes health measurements and logs related to interface selection and session failover.

That is one platform’s design, not a universal model. Platforms calculate health, evaluate thresholds, and treat failed or recovered paths differently. Ask your provider to explain your actual settings, in your installed software version, without hiding behind the phrase “dynamic path selection.”

Run at least three controlled conditions:

  • A hard loss of the primary path
  • A degraded path that crosses the policy’s loss, latency, or jitter threshold
  • Restoration of the primary path and the return to normal routing

Record what the platform measured, when it changed path, which sessions survived, which sessions restarted, and whether the alert reached the correct owner.

Test the application path the business needs

A green tunnel is not the finish line.

Choose a small set of business workflows that represent the site’s continuity requirement. Test them before, during, and after failover. Depending on the location, that may include:

  • Inbound and outbound voice
  • Payment or point-of-sale transactions
  • Warehouse or manufacturing transactions
  • Access to a cloud ERP or line-of-business application
  • Remote access and site-to-site VPN traffic
  • Identity, DNS, and cloud security services
  • Monitoring, logging, and remote management

Define what passes before the test starts. “The application came back” is too vague.

Measure detection time, path-change time, session impact, transaction recovery, available bandwidth, user-visible interruption, alert delivery, and failback behavior. If the backup link has less capacity, confirm that critical traffic gets priority and that nonessential traffic is limited as designed.

Public IP changes deserve their own check. A backup path may break a vendor allowlist, hosted service, VPN peer, voice registration, payment connection, or remote support tool even when ordinary web traffic works.

Make the provider show the current policy

The renewal proposal may list appliances, sites, bandwidth, licenses, and support. It may say very little about the policy now controlling production traffic.

Ask for a current export or provider-reviewed record covering:

Decision areaEvidence to require
Path inventoryEvery active transport, service ID, role, bandwidth, and site
Health detectionProbe targets, metrics, thresholds, timers, and failure conditions
Application steeringWhich applications prefer, avoid, or may use each path
Capacity treatmentPriority, shaping, and blocked traffic on constrained backup links
Security pathFirewall, SASE, DNS, inspection, and logging behavior on every path
FailbackWhen traffic returns and how unstable links are prevented from flapping
AlertingEvent, severity, destination, acknowledgement, and escalation owner
SupportWho changes policy, opens carrier cases, and leads recovery
Test evidenceDate, scope, result, gaps, corrections, and retest status

If the service includes SASE or managed security, verify that path changes do not bypass inspection or strand traffic behind the wrong policy. The SASE renewal policy ownership checklist can help separate who requests, approves, implements, tests, and rolls back those changes.

Do not accept a screenshot of the main dashboard as the complete record. It can show that links exist. It may not show why traffic chose one, what happens when no path meets the target, or whether the provider changed a default during the term.

Put the failover test into the renewal scope

Buyers often treat testing as an informal favor after the contract is signed. That is backwards when resilience is part of the value being renewed.

Define the work in writing:

  • Which sites and applications are in scope
  • Which failure and degradation conditions will be tested
  • Who approves the test window and business validation
  • Who monitors carriers, edge devices, applications, and security controls
  • What evidence the provider will deliver
  • Who owns corrections and carrier escalation
  • How quickly failed items must be retested
  • Whether testing is included, limited, or separately billable
  • What happens if the design cannot meet the agreed continuity requirement

Also clarify responsibility after future changes. A carrier swap, bandwidth upgrade, edge replacement, routing change, major firmware upgrade, or new critical application should trigger a review. The exact trigger list should match your environment, but “we test once at installation” is not enough.

Use four renewal outcomes

The test should lead to a decision, not another report nobody opens.

Renew as designed when current policy exports, monitoring, support ownership, and end-to-end test evidence match the business requirement.

Correct before renewal when the platform fits but the circuit inventory, thresholds, application rules, alert route, failback behavior, or support language needs repair.

Resize or redesign when backup capacity cannot carry critical work, physical paths share too much risk, security controls break during failover, or the operating model depends on manual work the business cannot perform quickly.

Run a competitive review when the provider cannot produce the current configuration record, will not perform a meaningful test, repeatedly leaves failures uncorrected, or cannot explain how the service behaves after material changes.

A failed test does not automatically mean the platform is wrong. Sometimes the circuit, application dependency, provider process, or customer approval path is the real issue. That is why testing before renewal is useful. You still have time to fix the right layer and negotiate the right obligation.

SD-WAN earns its place when it keeps important work moving through a real network problem. That value can decay quietly as circuits, applications, policies, people, and providers change.

Bring the current agreement, proposal, site and circuit inventory, topology, policy export, change records, monitoring history, incident tickets, and last test result into the renewal review. Choose the business workflows that must survive. Run the hard failure, degraded-path, and failback tests. Record the gaps. Retest the corrections.

If your SD-WAN or connectivity agreement is approaching renewal, request a Contract and Spend Risk Review. Catch Advisors will help you reconcile the circuits, service scope, support ownership, and continuity evidence before you renew, redesign, or compare providers.

Sources