CISA: IT Operations and Business Resilience

CISA Domain 4 asks whether technology operations can deliver reliable service and recover when ordinary controls are not enough. It joins daily operational disciplines—assets, capacity, incidents, change, patching, logging, service levels, and databases—with business-impact analysis, backup, continuity, and disaster recovery. That combination reflects an auditor’s real job: operational weakness and resilience weakness are often different stages of the same failure chain.

The current CISA assigns 26 percent to Information Systems Operations and Business Resilience. This domain asks auditors to evaluate whether operations and recovery practices support organizational objectives, using evidence from day-to-day service management as well as continuity and recovery controls.

The most useful mindset is evidence-first. A procedure may say that backups run, patches are tested, incidents are escalated, and recovery plans exist. The auditor asks whether records, metrics, tests, exceptions, and actual outcomes prove those statements.

Operational controls should support service objectives

Operations includes the technology components and processes that keep information systems useful: infrastructure, interfaces, scheduling, end-user computing, capacity, configuration, and service management. The auditor begins by understanding which services are important, who owns them, what service levels or business expectations apply, and which dependencies could prevent those expectations from being met.

Evidence may include inventories, runbooks, monitoring records, batch results, capacity reports, service-level reports, and exception logs. The purpose is not to demand documentation for its own sake. It is to establish whether operations are repeatable enough that performance and failures can be explained.

Asset and configuration knowledge reduce operational uncertainty

An organization cannot reliably operate or recover systems it cannot identify. Asset records, ownership, lifecycle state, supported versions, dependency mapping, and configuration baselines help teams understand what exists and what changed. Weak inventories create hidden risk because unsupported or forgotten components may sit outside patching, monitoring, backup, and recovery processes.

CISA audit work compares records with reality. Sampling physical, virtual, cloud, database, network, or application assets can reveal gaps between inventory systems and deployed environments. Configuration records should also support change investigation so an incident team can identify whether a deviation is approved, accidental, or malicious.

Availability and capacity should be managed before users complain

Capacity management looks ahead while availability management looks at whether services meet required uptime and reliability. Both depend on trend data, demand forecasts, thresholds, architecture limits, and business priorities. A service can be technically online yet unable to meet demand, which makes raw uptime an incomplete measure.

Auditors can review utilization trends, capacity plans, scaling decisions, single points of failure, backlog growth, and the relationship between service objectives and investment. Repeated emergency scaling or chronic saturation may show that the process is reactive even if individual incidents are eventually resolved.

Incident and problem management produce different evidence

Incident management restores service; problem management reduces recurrence by addressing causes or contributing conditions. Auditors should see clear prioritization, ownership, escalation, communications, resolution evidence, and handoff when recurring patterns require deeper analysis. High-severity incidents also create evidence for continuity and resilience reviews.

Metrics can include detection and restoration time, recurrence, reopen rates, backlog age, severity distribution, and corrective-action completion. Metrics need context: very fast closure may reflect efficient resolution or premature ticket closing. Sampling case records helps test whether performance statistics match the actual quality of decisions.

Change and patch management control operational risk

Changes introduce risk even when they are necessary. Effective processes classify changes, assess impact, define testing and approval, schedule implementation, prepare rollback, and review outcomes. Emergency changes should have a justified path that is faster without becoming an ungoverned bypass.

Patch management similarly balances exposure against service stability. The auditor examines inventory coverage, risk-based prioritization, testing, deployment evidence, exceptions, failed installations, and compensating controls. A headline patch percentage means little if critical systems or high-risk vulnerabilities consistently sit in the excluded remainder.

Logs and monitoring support both operations and assurance

Operational logs, security logs, job output, database events, infrastructure telemetry, and application monitoring help teams detect failure and explain what happened. Audit concerns include completeness, time synchronization, retention, access control, review ownership, and whether important alerts generate an appropriate response.

Log management is not the same as storing every event indefinitely. The organization should identify which evidence supports operations, security, compliance, and investigations, then retain and protect it accordingly. Missing logs can make root-cause analysis or incident reconstruction impossible even when other controls are strong.

Business-impact analysis connects technology to recovery priority

A BIA identifies critical activities, dependencies, tolerable disruption, and recovery priorities. From an IT perspective, it should translate business needs into recovery requirements for applications, data, infrastructure, people, facilities, and third parties. Recovery objectives that have no traceable business basis may be arbitrary or financially inefficient.

Auditors can test whether BIA assumptions are current and whether major business or technology changes trigger review. A new customer channel, outsourced service, cloud migration, acquisition, or critical data dependency can materially change recovery requirements even if the continuity document itself has not been updated.

Backups are useful only when restoration is proven

Backup success messages do not demonstrate recoverability. Organizations need defined scope, retention, protection from corruption or unauthorized deletion, off-site or logically separated copies where appropriate, and restoration testing. The auditor should trace important systems and data to a backup method and then look for evidence that recovery was tested.

Restoration tests should verify more than file presence. Applications may depend on identity, configuration, encryption keys, network paths, databases, and integration endpoints. A technically restored server that cannot rejoin the service does not satisfy the business recovery objective.

BCP and DR need exercises that expose real dependencies

Business continuity keeps priority activities operating while disaster recovery restores technology capability. Plans should define roles, invocation criteria, communications, alternate arrangements, dependencies, recovery sequences, and decision authority. Exercises test whether those assumptions survive contact with a realistic scenario.

The CISA certification expects auditors to evaluate whether continuity and recovery controls are designed and operating effectively. Tabletop exercises, technical failovers, restore tests, crisis communications, and post-exercise actions all provide evidence. Repeated findings that never close indicate a resilience program that documents risk without reducing it.

Domain 4 connects ordinary operations with extraordinary recovery. Asset knowledge, capacity, incidents, change, patching, logs, service levels, backups, BIA, continuity, and disaster recovery form one assurance chain: the organization must operate systems predictably and recover them when normal operating controls fail.

The CISA auditor’s role is to test that chain with evidence. Procedures explain intended behavior; records and exercises show whether the intended behavior occurs; and findings identify where operational or resilience weaknesses threaten business objectives.

Third-party dependencies belong in operational resilience even when the underlying infrastructure is not owned by the organization. Managed service providers, cloud platforms, telecom carriers, payment networks, software vendors, and external data services can all become critical recovery dependencies. Auditors should look for service commitments, escalation contacts, continuity obligations, concentration risk, and evidence that vendor recovery assumptions have been tested rather than copied into an internal plan without verification.

Job scheduling and automated production processes also deserve attention because they often fail quietly. Missed batch jobs, delayed data feeds, failed integrations, or incomplete reconciliations can damage business processes without taking a system offline. Controls can include job monitoring, dependency checks, restart procedures, exception queues, and reconciliation of expected versus completed work. Audit sampling should test both successful and failed runs to determine whether exceptions trigger reliable follow-up.

Recovery priorities can conflict when many services depend on the same constrained resource or specialist team. A plan that assigns every system a short recovery target may be internally impossible. Auditors should compare dependency maps, resource assumptions, alternate-site capacity, staffing, and recovery sequence so the plan reflects what can actually be restored. The best resilience evidence demonstrates coordinated priorities under realistic constraints, not a collection of optimistic targets maintained independently by application owners.

Database management is explicitly included in CISA Domain 4 because data services can fail even when surrounding infrastructure appears healthy. Auditors may examine backup/recovery, access administration, capacity, maintenance, integrity checks, job failures, replication, change control, and monitoring. The exact technology is less important than whether the organization can demonstrate consistent operation, protected data, recoverability, and appropriate ownership of database-level exceptions.

Service-level management converts operational performance into agreed expectations. SLAs and internal service objectives should define measurable outcomes, ownership, exclusions, and escalation when performance falls short. Auditors can compare reported results with underlying incident and monitoring data to detect optimistic reporting. A service that meets its average availability target while repeatedly failing during a critical business window may still be operationally inadequate.

Resilience improvement should continue after exercises and real incidents. Findings need owners, deadlines, priority, and evidence of closure. If a test repeatedly identifies the same missing dependency, capacity shortfall, or communication weakness, the program is rehearsing failure rather than reducing risk. CISA candidates should connect lessons learned back to change management, asset/configuration records, capacity planning, vendor oversight, backup design, and governance reporting.

Operational resilience also benefits from explicit recovery ownership. During a disruption, teams should know who can declare an incident or disaster, authorize failover, accept degraded operation, communicate externally, and decide when normal service is restored. Ambiguous decision authority can delay recovery even when technical procedures are sound.

  • img