Infrastructure Architecture / Field guide
Eight SD-WAN architecture decisions to make before vendor selection
SD-WAN selection should follow an architecture decision, not replace one. A defensible program begins by defining how applications, sites, transports, routing, security, resilience, operations, and migration must work together—then evaluates platforms against those requirements.
1. Application and traffic architecture
Document where applications live, who uses them, which paths they take today, and what performance or availability they require. SaaS, public cloud, private data centers, voice, real-time applications, partner connections, and remote access can create different path and inspection needs.
Avoid designing from average bandwidth alone. Latency, jitter, loss, session behavior, traffic direction, growth, peak periods, data sensitivity, and business criticality all influence the architecture.
2. Site tiers and connectivity patterns
Not every location needs the same design. Establish site tiers using business impact, user and device population, application dependency, operational access, and acceptable outage. Then define transport diversity, device resilience, local services, management, and support expectations for each tier.
A reference design should be repeatable without pretending every site is identical. Explicit exception criteria are part of the standard.
3. Underlay transport and provider strategy
SD-WAN does not remove underlay dependency. Define acceptable circuit types, access diversity, carrier diversity, bandwidth, service levels, addressing, handoff, demarcation, out-of-band options, and contract timing. Two circuits that share a last-mile path may not provide the resilience the diagram implies.
Record who owns carrier procurement, activation, testing, incident escalation, and chronic service-quality problems.
4. Routing, segmentation, and policy boundaries
Define how the overlay exchanges routes with campus, data center, cloud, internet edge, and partner networks. Address route ownership, summarization, redistribution, convergence, segmentation, shared services, overlapping address space, and coexistence during migration.
Path policy should connect to application and business requirements. It should also define what happens when telemetry is incomplete or every available path is degraded.
5. Security and inspection architecture
Decide where internet access, inspection, segmentation enforcement, remote access, cloud security, DNS controls, and logging belong. The answer may differ by site tier and application, but it should form one explainable trust model.
Security is most effective when it is built into topology, identity, routing, access, and operating ownership from the start—not bolted onto a connectivity design after the platform has been chosen.
6. Failure behavior and resilience
List the failures that matter: circuit, last mile, provider, device, power, DNS, authentication, controller, cloud gateway, security service, route exchange, and management-plane access. Define the expected user and application behavior, convergence target, degraded mode, alert, and recovery path for each.
A feature demonstration is not a resilience test. Acceptance should include controlled failure scenarios using the intended topology, policy, and operating process.
7. Operations, visibility, and lifecycle
Identify who monitors the service, interprets path analytics, approves policy, performs upgrades, administers access, maintains templates, handles exceptions, opens provider cases, updates documentation, and owns capacity and lifecycle decisions.
Consider how the platform integrates with identity, logging, ticketing, configuration governance, asset records, and incident response. Operational fit can distinguish two technically capable options.
8. Migration and vendor evaluation
Create a requirements matrix before product scoring. Evaluate architecture fit, supported integrations, routing behavior, security model, visibility, automation, licensing, support, hardware lifecycle, provider dependencies, deployment effort, and total operating implications.
The migration plan should cover pilots, representative site types, prerequisites, circuit readiness, coexistence, routing and security transitions, rollback, validation, documentation, training, and acceptance. A low-risk pilot is designed to expose assumptions before they reach the full rollout.
- Use cases and requirements traced to product evidence
- Lab or pilot tests for material routing, policy, and failure behavior
- Site-wave plan aligned to circuits, contracts, applications, and support capacity
- Explicit operational handoff and acceptance criteria
This guide provides general information for planning and evaluation. Exact technical, security, legal, and commercial requirements depend on the environment and agreed scope.
