Azure Monitor, Log Analytics, and Observability in Practice

Azure observability is easy to overcomplicate because the platform offers metrics, logs, traces, alerts, dashboards, workbooks, Application Insights, Log Analytics, data collection rules, and several specialized monitoring experiences. Production design becomes clearer when those services are treated as parts of one evidence system: detect that something is wrong, understand what changed, locate the failing dependency, and give operators enough context to decide what to do next.

The key is to design the telemetry around operational questions rather than around what each Azure service happens to expose by default. A large volume of logs is not the same as observability. Useful observability connects symptoms to causes.

Start With the Questions Operators Must Answer

Before configuring diagnostics, write down the decisions the platform needs to support. Can the team tell whether a service is available? Can it distinguish a dependency failure from an application failure? Can it see saturation before customers experience errors? Can it identify which deployment caused the change? Can it separate tenant-wide incidents from one noisy resource?

These questions lead naturally to a signal strategy. Metrics are excellent for fast detection and trend analysis. Logs provide event detail and investigative context. Traces show how a request moved across dependencies. Configuration and change data help connect symptoms to releases or policy changes. A good observability design uses each signal for what it does best.

The general principles are covered in observability fundamentals. Azure Monitor is the Microsoft-native implementation layer where those principles become operational.

Use Metrics for Fast Health and Capacity Signals

Azure Monitor metrics are time-series data optimized for fast analysis and alerting. They are the right choice for measurements such as CPU utilization, request rate, queue depth, failed requests, latency percentiles, memory pressure, storage transactions, or resource saturation.

A useful metric has a clear relationship to user or system health. Avoid alerting on every available measurement. A CPU spike may be harmless if latency and error rate stay healthy; a small increase in queue depth may be critical if it indicates work is no longer draining. Pair resource metrics with service-level signals so teams do not confuse internal activity with customer impact.

Baseline behavior by workload and time period. Static thresholds are simple but can create noise in environments with predictable peaks. Dynamic approaches can help, but they still need ownership and review.

Use Logs for Context, Investigation, and History

Logs are stored in Log Analytics workspaces and queried with KQL. They are where operators investigate the events behind an alert: which requests failed, which identity changed a resource, which dependency timed out, what the error payload contained, or how behavior changed after deployment.

Workspace design should begin with the simplest viable topology. Microsoft’s current guidance recommends starting with a single workspace for many environments, then adding more only when tenant, region, security-boundary, regulatory, cost, or ownership requirements justify the complexity. Multiple workspaces make cross-workspace query, access, retention, and operational ownership harder.

Retention and table plans should reflect investigative value. High-volume diagnostic tables may not deserve the same retention as security or audit data. The goal is to retain enough evidence for real incidents without paying to store data nobody uses.

Control Collection With Data Collection Rules

Data collection rules (DCRs) define how supported telemetry is collected, transformed, and routed. They are important because observability quality is shaped before data reaches a query. A DCR can standardize what fields are retained, filter noise, or transform records at ingestion.

Filtering can reduce cost, but it should be evidence-based. Removing a field because it looks unimportant may later make root-cause analysis impossible. Before filtering, identify which queries, alerts, dashboards, and investigations use that field. Treat changes to collection rules like production code: review them, test them, and monitor the effect on downstream detections.

For Application Insights data, workspace transformations can also be applied through a DCR. Propagation is not always instantaneous, so troubleshooting should account for configuration activation time before concluding that a transformation is broken.

Instrument Applications With OpenTelemetry and Application Insights

Infrastructure telemetry can show that a service is slow. Application telemetry can show why. Application Insights and OpenTelemetry instrumentation expose request traces, dependency calls, exceptions, custom events, and distributed context across services.

Trace correlation is especially valuable in microservice and AI architectures, where one user request may call APIs, databases, search systems, models, queues, and external tools. Without correlation identifiers or distributed tracing, teams end up searching several systems manually and guessing which events belong together.

Observability should also connect releases to behavior. Deployment version, model version, feature flag, region, tenant, and other operational dimensions can make the difference between a general alert and an actionable diagnosis.

Design Alerts Around Actionability

An alert should tell someone that action may be required. If the expected response is always “ignore it,” the alert is not useful. Start with service symptoms, failure thresholds, and capacity conditions that matter to customers or operators, then define the owner and the first investigative step.

Use action groups to route notifications or trigger automation, but avoid building a chain of automatic reactions before the alert quality is proven. A noisy alert tied to a remediation runbook can create automated instability.

Alert rules need lifecycle management. Review whether they still fire, whether thresholds match the current scale, whether the owner still exists, and whether the condition is now covered by a better health model or SLO.

Use Workbooks and Dashboards to Support Decisions

Dashboards and workbooks should answer operational questions, not display every available graph. A service-owner view may need availability, error rate, latency, saturation, top dependencies, recent deployments, and active incidents. A platform-team view may emphasize ingestion, workspace health, agents, DCRs, and data cost.

Log Analytics Workspace Insights can help administrators understand workspace usage, performance, agent health, query activity, DCRs, and configuration changes. Those operational views are valuable because the monitoring platform itself can fail or drift. If a connector stops sending data, the absence of alerts is not necessarily good news.

Make Cost Part of Observability Architecture

Telemetry cost is an architecture concern because collection, ingestion, retention, query, and export choices can grow faster than the workloads being monitored. Cost governance begins with knowing which tables drive ingestion, which diagnostics create duplicate data, and which fields or events are actually useful.

Do not optimize only by deleting data. Sampling, transformation, retention tiers, table plans, better dimensions, and reduced duplication can preserve useful evidence while controlling spend. The broader reliability, security, performance, cost, and operations apply here: cost, reliability, security, performance, and operations are coupled.

Troubleshoot the Observability System Itself

When expected telemetry is missing, check the chain in order: is the source producing data, is the agent or diagnostic setting configured, is the DCR associated correctly, is the destination healthy, is the data arriving under the expected table/schema, and does the querying identity have permission to see it?

For alerts, separate “data not present” from “query did not match” from “action group did not notify.” For Application Insights, separate instrumentation failure from sampling or filtering. For multi-workspace designs, verify that queries are pointed at the intended workspace and tenant.

A mature Azure Monitor environment is observable itself. Teams know ingestion health, workspace changes, query performance, agent state, and alert delivery health—not just application health.

Multi-region systems need an observability design that survives the same failures as the application. If every dashboard, alert, or query path depends on one region or one workspace that becomes unavailable, operators can lose visibility during the outage they most need to understand. Critical services should consider how telemetry, synthetic tests, status views, and escalation function during regional impairment.

Deployment markers should be captured automatically. A chart that shows latency rising at 14:07 is much more useful when the same timeline shows that version 4.12 was deployed at 14:05 and a configuration flag changed at 14:06. Release systems can annotate monitoring platforms or emit custom events so operators can correlate changes without manually comparing deployment records.

Alert ownership should follow service ownership. Platform teams can own shared infrastructure alerts, while application teams own business and code-level signals. Routing every alert to one central operations queue creates handoff delays and encourages alert fatigue. Define primary owner, escalation target, and after-hours expectations for each alert class.

Post-incident review should include the observability system itself. Ask whether the first alert represented user impact, whether responders had the right logs, whether traces were complete, whether dashboards accelerated or delayed diagnosis, and whether retention preserved necessary evidence. Every significant incident is an opportunity to improve what the system will reveal next time.

Telemetry design also needs a clear data contract. Teams should know which resource dimensions, correlation IDs, tenant IDs, operation names, deployment versions, and business identifiers are expected in each signal. When those fields change silently, dashboards and alerts can remain syntactically valid while becoming semantically useless. Treat schema changes as production changes, especially for custom logs and application telemetry that feed operational or security decisions.

Sampling deserves careful treatment. Traces are often the most expensive high-cardinality signal, so sampling can control volume, but careless sampling can remove the rare failures operators most need to inspect. Keep error and high-latency scenarios at higher fidelity, preserve trace correlation where possible, and understand whether sampling decisions happen in the SDK, collector, ingestion path, or query layer. An investigation is difficult when one service kept the trace while another dropped the matching span.

Alert design should include failure of the monitoring pipeline itself. Heartbeats, data freshness checks, synthetic availability tests, and dead-man alerts can detect when expected telemetry stops arriving. This is especially important for security and regulated workloads because a quiet dashboard may mean the environment is healthy, or it may mean the data source is broken. Distinguish “zero events” from “no data received.”

Access is another architecture decision. Log Analytics data can contain identities, IP addresses, request payload fragments, business identifiers, and application errors that expose sensitive context. Use RBAC, table-level controls where applicable, and separate workspaces only when the requirement justifies the operational complexity. Query access should follow least privilege, and export or diagnostic pipelines should not create uncontrolled copies of the same sensitive telemetry.

During incident review, connect telemetry cost to usefulness. A table that consumes large ingestion volume but is never queried may be a candidate for filtering or shorter retention. A small audit table that proves who changed a critical resource may deserve long retention. Cost optimization is not “store less”; it is “pay for evidence in proportion to its operational value.”

Telemetry freshness should be monitored explicitly. A quiet dashboard can mean a healthy system or a broken collection pipeline. Track the last-seen time for important tables, agent heartbeat, diagnostic-setting state, and synthetic-test execution. Operators should be alerted when evidence disappears, especially for sources used by security, compliance, or high-severity service alerts.

Deployment markers make time-series data far more useful. Emit release version, configuration change, and feature-flag events into the same observability timeline used for latency and failures. When an incident begins minutes after a change, responders can test that hypothesis immediately instead of manually comparing deployment systems and monitoring timestamps.

Capacity analysis is a slower but equally important observability use case. Long-term metrics can reveal saturation thresholds, seasonal peaks, query growth, or ingestion trends before service objectives are violated. Capacity dashboards should be separated from paging alerts so teams can plan upgrades without turning every gradual trend into an operational emergency.

Multi-region services should test whether monitoring remains usable during the failures it is intended to diagnose. If all alerts, workbooks, or synthetic checks depend on one region, a regional outage can remove visibility at the same time the application fails. Critical systems need an observability path that survives their credible failure modes.

Post-incident review should ask whether the monitoring system helped. Was the first alert aligned with user impact? Were the right logs retained? Did distributed traces cross the failing dependency? Did the dashboard shorten diagnosis or create distraction? Incidents are the best source of requirements for the next iteration of observability design.

Resilience of the monitoring plane should be reviewed alongside workload resilience. If operators depend on one workspace, region, dashboard, or notification channel during an outage, verify that this dependency remains available in the failure scenarios the workload is designed to survive. Cross-region recovery plans should state where telemetry will be queried, how alerts continue, and whether diagnostic settings or workspace replication need special handling.

What Production Observability Looks Like

A production-ready Azure observability stack has deliberate signal ownership. Metrics detect meaningful health changes. Logs preserve enough context for investigation. Traces connect distributed requests. DCRs are governed as code. Alerts are actionable and routed to owners. Dashboards answer operational questions. Workspace topology is as simple as requirements allow.

For people working toward Azure architecture skills, the existing AZ-305 architecture skills is a natural destination for broader architecture preparation. In live systems, the real test is whether engineers can move from “something is wrong” to “this is the failing dependency, this change caused it, and this is the safe next action” without stitching evidence together manually.

  • img