Private Link and Private Endpoint Architecture in Production

Azure Private Link is often described as a way to access a platform service over a private IP, but that summary hides the operational work. A private endpoint adds a network interface in a virtual network and connects it privately to a supported service. The application still depends on identity, authorization, routing, firewall policy, and—most importantly—correct DNS. Creating the endpoint is only one step in the architecture.

Production designs fail when teams treat Private Link as a checkbox instead of a name-resolution and access-path design. The reliable approach is to model who needs to reach the service, from which networks, using which FQDN, through which resolver path, under which identity, and with what public-access policy.

Understand What a Private Endpoint Actually Changes

A private endpoint creates a private IP address in your virtual network for a specific supported service or subresource. Traffic to that private IP uses the Microsoft backbone rather than reaching the service over its public endpoint. The endpoint has a network interface and lifecycle of its own, and many services require you to choose the exact subresource being exposed.

That does not automatically mean the service’s public endpoint is disabled. Public access is a separate service configuration. A secure design therefore has two questions: can the workload reach the private endpoint, and is the public access path still allowed? Teams that answer only the first question often believe they have created isolation when they have only added an additional path.

Private Link and private endpoints for SC-500 establish the security model. Production design adds DNS, operations, dependency mapping, and failure isolation to that foundation.

Design DNS Before You Deploy the Endpoint

The service FQDN is usually the application contract. Applications should keep using the service name rather than being rewritten to a raw private IP. DNS is what makes that possible by resolving the service’s private-link name to the private endpoint address for workloads that should use the private path.

Azure Private DNS zones are the common mechanism inside Azure. The right zone depends on the service. Link the zone to the virtual networks that need resolution, and make sure records are created for the correct resource. When multiple environments or endpoints for similar services exist, avoid blindly sharing zones in ways that overwrite or remove required records.

Microsoft explicitly warns that DNS design is critical. A private endpoint can be healthy while the application still reaches a public address—or cannot resolve the service at all—because the FQDN is being answered by the wrong DNS authority.

Plan Hybrid and Hub-Spoke Name Resolution Explicitly

Hybrid environments are where Private Link DNS designs become difficult. On-premises clients may query corporate DNS servers that know nothing about Azure Private DNS. Spoke virtual networks may use a hub resolver. Multiple regions may need resilient resolution. The design has to describe the forwarding path end to end.

Azure DNS Private Resolver can provide managed inbound and outbound resolution without maintaining custom DNS virtual machines, but it still requires correct rules and network reachability. On-premises DNS can forward relevant private-link zones to an inbound endpoint, while Azure workloads can use outbound rules when they must resolve private namespaces elsewhere.

Document the exact FQDN, which DNS server answers first, which conditional forwarder is used, and which private zone owns the record. During an outage, this map is far more valuable than a diagram that only shows the private endpoint resource.

Coordinate Routing, Firewalls, and Network Security

A correct private DNS answer is necessary but not sufficient. The client still needs a route to the private IP and a path that is not blocked by NSGs, firewalls, network virtual appliances, or host-level policy. Central egress inspection can also create asymmetric paths if the return route does not match the forward route.

Private endpoints are commonly deployed into dedicated subnets so policy and ownership are clear. Be deliberate about NSG and route-table behavior for those subnets, and understand the supported network-policy settings for your scenario. Do not copy controls from a normal application subnet without checking how the service and endpoint behave.

The general principles in Azure network security group design still apply: traffic paths, allowed flows, logging, and ownership should be explicit.

Use Identity and Authorization Alongside Network Isolation

Private Link limits the network path; it does not decide which application identity may read a secret, query a database, download a blob, or invoke an AI model. Authorization remains a service-level concern. Use managed identities, Microsoft Entra authentication, RBAC, database permissions, or other supported controls to enforce least privilege.

This is especially important in shared virtual networks. A workload that can reach a private endpoint should not automatically gain data access. The network proves where the request came from, while the identity and service authorization determine what the caller is allowed to do.

Think of the private endpoint as one layer in a defense-in-depth design, not as the security boundary by itself.

Decide How Public Access Will Be Treated

Some teams want a fully private service and should disable or tightly restrict public network access after private connectivity is validated. Others need both public and private paths during a transition or for different client populations. The design should state which state is intended instead of leaving it to default behavior.

When public access remains available, ensure the application and DNS design do not accidentally bypass the private path. When it is disabled, validate all management, monitoring, automation, and disaster-recovery dependencies that may have used the public endpoint previously.

Migration plans should include an observation period where connection logs and DNS behavior can be reviewed before the public path is removed.

Troubleshoot From Name to Route to Authorization

A reliable troubleshooting sequence starts with DNS. Resolve the service FQDN from the failing client and confirm that it maps to the expected private IP. If it does not, inspect DNS suffixes, forwarding rules, private-zone links, and record registration.

If DNS is correct, test reachability to the private IP and validate routes, NSGs, firewalls, and peering. Then move up the stack to TLS, service configuration, identity, and authorization. A 403 response is a different class of problem from a timeout; an NXDOMAIN response is different again.

Use the service’s own diagnostic logs where available. Connection approval state, endpoint health, NIC configuration, and public-network settings can narrow the problem quickly.

Design for Multiple Regions and Recovery

Private connectivity needs a recovery story. If the service fails over to another region or a secondary instance becomes active, clients need DNS and routing that can follow the change. Some services use different regional FQDNs or private endpoints; others require application-level failover logic.

Do not assume that replicating the service automatically replicates the private access path. Recovery runbooks should verify endpoint creation or readiness, DNS records, zone links, resolver health, and security policy in the recovery region.

This is where broader cloud disaster-recovery design becomes directly relevant: network and name-resolution dependencies are part of RTO, not separate infrastructure details.

Capacity and address planning matter because each private endpoint consumes an IP address in its subnet. Large estates that create endpoints for many resource instances and subresources can exhaust address space unexpectedly. Reserve subnet capacity with realistic growth assumptions and avoid mixing unrelated endpoint estates into tiny ranges that require disruptive readdressing later.

Service lifecycle events also affect DNS. Replacing a resource, moving to a new account, or changing regions can produce a new private endpoint and IP. Automation should update records and validate consumers before the old endpoint is removed. Stale private DNS records are particularly confusing because name resolution appears private while traffic points to an address that no longer serves the intended resource.

Build and deployment agents are a frequently missed client class. A CI/CD runner may need to fetch packages, deploy to a private service, or validate an endpoint from a network that lacks private DNS or routing. Teams sometimes open public access temporarily for pipelines, undermining the architecture. Instead, place runners on supported private networks or provide controlled connectivity that follows the same DNS model as production clients.

Recovery planning should include the private connectivity layer. If a service fails over to another region, will the application use a different private endpoint, a new DNS record, or a global service abstraction? Regional DR testing must validate private DNS and routing in the recovery region, not only the data service. An application that has replicated data but no working private name resolution has not achieved recoverability.

Service-specific subresources are another source of failure. Storage accounts, databases, AI services, and other platforms can expose separate endpoints for different capabilities. Creating a private endpoint for one subresource does not guarantee that an application using another hostname is covered. Inventory the FQDNs the application actually calls, including management, data, ingestion, and auxiliary endpoints, then map each one to the required private-link configuration.

DNS split-horizon behavior should also be tested from every network class. A laptop on the corporate network, a workload in a spoke VNet, a managed build agent, and a disaster-recovery environment may all query different resolvers even when they use the same FQDN. A design that works from one jump box is not validated. Build a small matrix of client location, DNS server, expected answer, route, and service authorization, and keep it with the runbook.

Private endpoints can complicate platform operations when deployment tools or managed services are outside the trusted network path. Before disabling public access, test CI/CD agents, Azure-hosted automation, monitoring services, backup tools, scanners, data-integration runtimes, and vendor integrations. If one of those requires network access, decide whether to move it into the private network, provide an approved path, or accept a controlled public exception.

Resource ownership matters because DNS zones, VNets, private endpoints, and the target service are often owned by different teams. Decide which team approves endpoint connections, which team owns private DNS records, and who is responsible when a service is deleted or recreated. Orphaned DNS records can send traffic to dead addresses, while a recreated endpoint may receive a new private IP and leave cached or manually maintained records stale.

Observability should include DNS query failures, connection failures, endpoint state, firewall denies, and service-side authentication errors. When these are collected centrally, a team can distinguish an access-control problem from a network-path problem quickly. Without that evidence, Private Link incidents become long cross-team calls where each group proves only that its own component is healthy.

Address planning becomes important at scale because each private endpoint consumes a private IP. A small subnet that works for a pilot can become a hard limit when applications add multiple resource instances and subresources. Reserve realistic growth capacity and avoid placing unrelated endpoint estates into ranges that would require disruptive readdressing later.

Build agents and automation runners are often forgotten clients. A deployment pipeline may need to reach a private storage account, Key Vault, registry, or database from a network that lacks the required DNS and routes. Opening public access temporarily for pipelines undermines the architecture. Place runners on supported private networks or design a controlled connectivity path from the beginning.

Regional recovery needs a private-connectivity plan as well as a data plan. If an application fails over to another region, determine which private endpoint, DNS record, and route will represent the recovered service. Test the private name-resolution path during disaster-recovery drills. Replicated data without reachable private endpoints is not a recoverable application.

Lifecycle automation should clean up as reliably as it creates. Deleting a resource can leave stale private DNS records, unused endpoint network interfaces, or shared-zone links. Stale records are especially confusing because DNS still returns a private address that no longer serves the intended target. Make deletion and replacement part of the module contract.

Enterprise policy can reinforce the design by auditing public network exposure or requiring approved private-connectivity patterns for selected services. Introduce policy with a supported implementation path and an exception process. Denying deployments without reusable modules or clear DNS guidance simply shifts the problem from insecure architecture to blocked delivery.

Private-endpoint lifecycle should be automated where possible, but automation must include DNS and approval state rather than creating only the network interface. A deployment is not complete until the service connection is approved, the expected private record exists, and a probe from the intended client network resolves and connects successfully. These post-deployment tests catch the most common gap between infrastructure state and application usability.

A Production Private Link Checklist

Before declaring a Private Link design complete, confirm the target subresource, endpoint approval state, private IP, FQDN, private DNS zone, VNet links, hybrid forwarding, routing, NSGs/firewalls, service authorization, public-access policy, monitoring, and recovery path. Test from each client class that will actually use the service.

For Azure administrators and architects, the AZ-104 administration skills and AZ-305 architecture skills are useful broader paths. In production, the decisive skill is being able to explain exactly why a hostname resolves to a particular private IP and exactly what happens to the request after that.

  • img