Transit Gateway Design in Production

AWS Transit Gateway turns a collection of VPC, VPN, and Direct Connect connectivity requirements into a regional routing hub. That can simplify large networks, but it also creates a new control plane whose route tables, attachment associations, propagation, and ownership determine which networks can communicate. A single Transit Gateway does not automatically create a safe hub-and-spoke architecture; it creates the mechanisms from which one can be built.

AWS publishes specific Transit Gateway design practices around attachment subnets, BGP, route-table design, and resilience. Production architecture should use those mechanics to create understandable routing domains across AWS rather than treating the Transit Gateway as one giant shared router.

Begin with routing domains, not attachments

List the groups of networks that should be able to communicate before creating Transit Gateway route tables. Production workloads may need shared services but not development VPCs. Inspection VPCs may need to see traffic between domains. On-premises networks may have access to only selected environments. These desired communication sets should drive route-table design.

A common failure is to attach every VPC, enable broad propagation, and only later try to reconstruct segmentation. That produces a large implicit trust zone. Transit Gateway route tables are easier to reason about when each association has a clear purpose and propagation is enabled only where the receiving domain should learn the routes.

Separate association from propagation in your mental model

An attachment is associated with one Transit Gateway route table, which is used to route traffic arriving from that attachment. An attachment can propagate routes to one or more route tables, depending on configuration. These are different controls. Misunderstanding them leads to “the route is visible but traffic still does not flow” or the reverse, where a network learns a route it was not meant to receive.

Document both sides. For each attachment, record the table it uses for outbound decisions and the tables into which its routes are advertised. A matrix is often clearer than a network diagram alone. The broader cloud networking concepts remain useful, but Transit Gateway adds a second routing layer beyond each VPC subnet route table.

Use dedicated attachment subnets and keep the data path explicit

AWS guidance recommends dedicated small subnets for Transit Gateway VPC attachments. Those subnets are not ordinary application subnets; they provide the attachment ENIs used by Transit Gateway in each selected Availability Zone. Keeping them separate makes routing and network ACL behavior easier to control and inspect.

Application subnets need routes that point to the Transit Gateway for remote prefixes, and the Transit Gateway route table needs a route toward the destination attachment. The return path needs equivalent reachability. If centralized inspection is inserted, every additional hop must be represented deliberately. Many production failures are not Transit Gateway failures at all—they are missing or asymmetric VPC routes around the attachment.

Design segmentation before enabling route propagation

Propagation is convenient because it reduces manual route maintenance, but convenience can widen trust. A shared-services route table may reasonably learn many workload prefixes. A production workload route table should not necessarily learn every sandbox or acquisition network. Where a strict boundary matters, use separate route tables and explicit propagation choices.

Blackhole routes can be used for specific containment patterns, but the architecture should not rely on a long list of exceptions to compensate for an over-broad base design. Prefer positive segmentation in which route domains are narrow by default. This makes later audits and incident containment much easier.

Create representative traffic tests for every routing domain, including paths that must remain blocked. Positive connectivity tests alone do not prove segmentation. A production validation suite should confirm that development cannot reach production where prohibited, that shared services are reachable from intended domains, and that inspection is not bypassed by a newly propagated prefix.

When an acquisition or temporary environment is connected, avoid placing it directly into the same route domain as trusted workloads. Use a dedicated route table, narrow advertised prefixes, and an explicit migration plan. Transit Gateway makes connectivity easy enough that temporary exceptions can otherwise become permanent trust relationships.

Integrate centralized inspection without creating asymmetry

Transit Gateway is frequently paired with an inspection VPC containing AWS Network Firewall, Gateway Load Balancer appliances, or third-party security appliances. Stateful inspection requires the forward and return flows to traverse compatible paths. Transit Gateway appliance mode can preserve Availability Zone affinity for traffic through an appliance VPC where the design requires it.

Build and test the complete route chain: source subnet to Transit Gateway, Transit Gateway to inspection, inspection VPC route back to Transit Gateway, Transit Gateway to destination, then the reverse direction. Draw the route tables, not only the boxes. Scale events and Availability Zone failures should not silently bypass inspection or create asymmetric return traffic.

Availability Zone placement matters in inspection architecture. Transit Gateway attachments are created in selected subnets, and stateful appliances may need zone-aware routing to preserve symmetric flows. If traffic enters through an attachment in one zone and the return path is attracted through another appliance path, sessions can fail even when every individual route looks valid. Test zone loss and scaling behavior explicitly.

Inspection bypass deserves a negative test. Add a representative route or attachment in a non-production environment and verify that the architecture still forces the intended inspection path. This catches designs where one direct propagation or static route can accidentally create an uninspected shortcut around the security VPC.

Treat hybrid BGP design as part of Transit Gateway architecture

VPN and Direct Connect integrations add dynamic routing and external administrative domains. AWS recommends unique BGP Autonomous System Numbers where appropriate and careful prefix planning. The goal is to make path preference and failover understandable rather than relying on accidental route selection.

Know which routes originate on-premises, which are propagated through Transit Gateway, and which AWS prefixes are advertised back. Summarization can reduce route scale but may also advertise reachability more broadly than intended. Hybrid failover should be tested with real withdrawal and reconvergence behavior, not inferred from a static architecture diagram.

Use inter-Region peering for deliberate regional connectivity

Transit Gateway is a regional resource. Inter-Region peering can connect regional Transit Gateways, but it should not turn every Region into one flat network. Decide which prefixes must cross Regions, how failure is detected, and whether the application can tolerate the additional dependency.

AWS guidance commonly favors a Transit Gateway per Region when the organization needs regional routing and disaster-recovery capability. Regional independence can be more valuable than maximal connectivity. Architects working with SAP-C02 should remember the exam transitions to SAP-C03 on November 17, 2026, while the multi-Region design principles remain applicable.

Do not assume that inter-Region peering should carry application replication, user traffic, management, and backup flows through one undifferentiated route domain. Separate traffic classes where their security or recovery requirements differ, and verify that a regional incident cannot accidentally reroute all control traffic through the region that is impaired.

Plan route scale, quotas, and ownership early

As VPC and on-premises networks grow, route counts, attachment counts, and operational changes become platform concerns. Track quotas before they become incident triggers. Establish who may create attachments, change route tables, enable propagation, and advertise prefixes. A central network team can own the Transit Gateway without owning every application subnet.

Automate repeatable attachment patterns and require metadata that identifies owner, environment, and routing domain. That helps prevent orphan attachments and makes cost allocation possible. Central networking becomes sustainable when teams can request a known connectivity product rather than negotiate custom routes every time.

Cost is also part of hub design. Attachment-hour charges and data-processing charges can make unexpected east-west paths visible on the bill. Central network teams should expose enough usage data for application owners to understand when architecture choices create avoidable transit. Cost anomalies can also reveal routing changes that were not planned.

Configuration drift is especially dangerous in a hub. Manage Transit Gateway route tables and attachments through reviewed automation where possible, but preserve an emergency operating path for incidents. Every automated change should have a clear rollback because a single propagated-route error can affect many otherwise healthy VPCs at once.

Emergency route changes need a post-incident cleanup step. Temporary static routes, disabled propagation, or bypass paths can remain after service is restored and quietly weaken segmentation. Record every emergency network change, assign an owner, and verify the intended steady-state route tables after the incident closes.

Troubleshoot by verifying every route domain in sequence

When traffic fails, confirm the source subnet route first, then the Transit Gateway attachment and association, the Transit Gateway route selected for the destination, any propagation or static route involved, the destination VPC route table, security controls, and the return path. Use flow logs and appliance telemetry to locate the break.

Do not assume that seeing a propagated route means the attachment is using the table that contains it. Association and propagation are separate. The AWS VPC routing layer still matters on both ends, which is why Transit Gateway troubleshooting requires a layered route model.

The value of Transit Gateway is not merely that many networks can connect. It is that a platform team can define which transitive paths exist and manage those paths consistently. The architecture should make trust domains, inspection paths, hybrid advertisements, regional boundaries, and failure behavior visible.

ANS-C01 remains current in October 2026 but is scheduled to retire on December 31, 2026. Whether used for certification preparation or production design, Transit Gateway skill should be measured by the ability to predict packet paths and blast radius—not by the ability to create another attachment.

  • img