ISACA CISA Deep Dive: Systems acquisition and implementation and Operations and resilience in Real-World Scenarios
The current CISA blueprint assigns 12 percent to Information Systems Acquisition, Development and Implementation and 26 percent to Information Systems Operations and Business Resilience. Studying them separately is necessary, but real systems cross the boundary at go-live. A project can satisfy its schedule and budget yet fail operationally because monitoring, capacity, ownership, backup, support, or recovery were never made ready. Conversely, an excellent operations team can spend years compensating for control weaknesses that should have been addressed during acquisition and implementation.
For exam scenarios, the useful question is not merely “Is this a project problem or an operations problem?” Ask where the control should have been established, what evidence exists now, and which audit objective is being tested. The auditor is evaluating whether management’s lifecycle controls protect business outcomes. That requires understanding business cases, project governance, development methods, control requirements, testing, migration, configuration and release management, post-implementation review, service operations, change, incident and problem management, availability, capacity, backups, business continuity, and disaster recovery.
The scenarios below focus on decision points and evidence. They deliberately avoid a single preferred technology because CISA is not a product exam. The same control objective can be achieved through different tools if responsibilities, approvals, evidence, and outcomes are sound.
A manufacturer wants a new cloud-based planning platform. The business case emphasizes subscription savings and faster deployment. Security, privacy, integration, data-retention, resilience, and exit requirements are deferred until contract negotiation. By then, changing the architecture would delay the launch and reduce the expected savings.
The control failure began before implementation. A complete business case should consider feasibility, costs, benefits, risks, dependencies, compliance obligations, resource needs, and alternatives. Audit readiness means checking whether decision makers received enough information to understand the full lifecycle commitment. If important requirements were intentionally deferred, the auditor should determine whether the residual risk was identified, accepted by appropriate authority, and reflected in contract and project plans.
Do not confuse “the project has executive approval” with “the project has adequate governance.” Approval is only as good as the information supporting it. The auditor may inspect the business case, steering-committee records, risk assessments, architecture decisions, vendor due diligence, requirements traceability, and budget assumptions. If exit costs or resilience dependencies were omitted, management may have approved an incomplete economic and risk picture.
A tempting response is to require more security controls immediately. That may be necessary, but the deeper issue is governance over acquisition. The durable correction is to embed relevant control and operational requirements before solution selection creates expensive lock-in.
A bank uses agile teams and continuous delivery. Auditors ask for a traditional phase-gate document set and conclude that controls are missing because the teams cannot produce a single signed design package. The teams, however, maintain approved user stories, architecture decisions, code review records, automated tests, pipeline logs, deployment approvals, and traceable production releases.
The auditor should evaluate control objectives, not insist on a historical methodology. Current development approaches can provide strong evidence if requirements, approvals, segregation, testing, traceability, and monitoring are built into the workflow. The key questions are whether business and control requirements are defined, changes are reviewed by appropriate people, code and configuration are tested, production releases are authorized, artifacts are traceable, and exceptions are managed.
Automation changes the evidence source. A pipeline may prove that only signed artifacts from an approved branch can deploy after tests pass. That can be stronger than a manually signed checklist if the pipeline is itself protected. Therefore the auditor should assess repository administration, pipeline change rights, service identities, secret handling, test bypass controls, artifact integrity, logging, and emergency override paths.
A readiness trap is assuming that DevOps eliminates segregation of duties. It changes how segregation can be achieved. Peer review, protected branches, independent approvals for high-risk releases, restricted production credentials, and immutable pipeline evidence can separate incompatible actions without preventing frequent delivery.
A healthcare organization buys a commercial SaaS platform. Because the vendor develops and hosts the software, the project team removes secure-development and infrastructure tasks from its plan. It also assumes the vendor is responsible for data classification, user access, configuration, retention, backup objectives, incident communication, and legal hold.
Acquisition changes control ownership but does not erase customer responsibilities. The auditor should identify which controls are provided by the vendor, which are configured by the customer, and which require coordinated activity. Contract terms, service descriptions, assurance reports, customer configuration guides, data-processing terms, incident commitments, recovery objectives, and exit provisions can help define the boundary.
Configuration deserves special attention. SaaS can reduce server administration while increasing dependence on tenant settings: identity federation, role design, sharing, data export, retention, audit logging, integrations, API credentials, and administrative privileges. A secure vendor platform can still be operated insecurely by the customer.
Third-party assurance evidence must be read for scope. Does it cover the actual service and period? Are subservice organizations excluded? Are there complementary customer controls that management must perform? Were exceptions identified? A report can reduce audit effort, but only after the auditor determines what conclusion it can legitimately support.
A retailer’s new order platform passes user acceptance testing. The business wants immediate go-live. Operations has not finalized monitoring thresholds, on-call escalation, capacity baselines, backup restoration, support documentation, privileged-access reviews, or vendor contact paths. Management argues that these items can be completed after launch because the application works.
User acceptance answers an important question: does the system meet business needs from the user’s perspective? It does not prove operational readiness. The auditor should evaluate whether implementation criteria include technical, security, support, resilience, data, and control requirements. Evidence might include readiness checklists, operational acceptance, runbooks, monitoring tests, support training, access reviews, backup and recovery results, capacity tests, unresolved-defect assessments, and formal go-live approval.
Risk-based judgment matters. Not every low-risk system needs the same ceremony as a payment platform, but critical services should not depend on undocumented assumptions. If management accepts a known gap, the decision should identify the risk, owner, remediation date, compensating control, and approval authority.
The audit lesson is timing. Operational controls are cheapest to establish while architecture and project resources are still available. Deferring them until after release often creates a period in which the service is technically live but not governable.
A company migrates ten million customer records to a new platform. Source and target record counts match exactly, so the project declares the conversion successful. Weeks later, the business discovers that date formats were transformed incorrectly, inactive accounts were reactivated, several regulatory flags were truncated, and historical relationships between records were lost.
Reconciliation must test more than row counts. A robust conversion plan identifies source populations, field mappings, transformation rules, validation criteria, rejected records, duplicates, referential relationships, totals, critical attributes, and exception handling. The auditor should determine whether management tested completeness, accuracy, validity, and authorization at a level appropriate to the risk.
Cutover also needs control. Who can alter source or target data during migration? When is the source frozen? How are late transactions handled? What rollback conditions exist? Who approves final reconciliation? Are conversion scripts version controlled and independently reviewed? If sensitive data is copied into staging locations, are those locations protected and removed afterward?
Sampling can be useful, but critical fields may require full-population analytics. If every regulatory classification must survive migration, testing only a small random sample may not address the risk. The auditor’s method should follow materiality and the nature of the control objective.
A software company has a disciplined release pipeline for routine changes. During outages, senior engineers log directly into production, modify configuration, and document the change later if time permits. Emergency access is widely shared because management prioritizes service restoration.
Emergency change capability can be legitimate, but urgency does not remove accountability. A controlled process should define who may invoke emergency access, what approvals are practical, how actions are logged, how configuration is captured, when normal testing can be deferred, how rollback is handled, and how retrospective review occurs. The auditor should test actual emergencies, not merely read the policy.
Configuration management is central. If engineers change production directly, the organization needs a way to reconcile the live state with the approved source of truth. Otherwise the next automated deployment may overwrite the emergency fix or reproduce an outdated configuration. Drift detection, post-event code updates, configuration baselines, and review evidence can close the loop.
A common distractor is to ban all emergency change. That can create unacceptable outage risk. The stronger control balances continuity with traceability and review, keeping emergency paths narrow and observable.
An enterprise completes an expensive workflow automation project. The implementation is declared successful because it launched on schedule. Six months later, many users still perform manual workarounds, expected headcount savings never occurred, and control exceptions increased because approval rules were configured too broadly.
A post-implementation review should evaluate whether the delivered system achieved business objectives, whether costs and benefits align with the business case, whether controls operate as intended, whether unresolved issues remain, whether operational ownership is stable, and what lessons should inform future projects. It is not merely a retrospective meeting for the project team.
Audit evidence can combine performance metrics, user adoption, incident records, control exceptions, benefit tracking, operating costs, support volume, and management review. If expected benefits are not measurable, that itself may reveal weakness in the original business case.
The important distinction is between project completion and outcome realization. A successful deployment can still be a failed investment or an ineffective control environment.
A financial system runs nightly settlement jobs. The scheduler reports success, yet a downstream interface occasionally rejects transactions because of malformed data. Rejected records sit in an error queue that no team owns. Month-end reconciliation finally reveals missing entries.
Operations assurance should follow the full data flow. Job completion is not enough. The auditor should evaluate scheduling, dependencies, interface controls, error handling, reconciliation, monitoring, ownership, retry logic, and escalation. Automated controls may include record counts, hash totals, sequence checks, duplicate detection, acknowledgments, and exception alerts.
The evidence question is whether management can identify incomplete or failed processing promptly enough to protect the business objective. A green scheduler dashboard may be reliable evidence that the job ran, but irrelevant to whether every transaction reached the destination accurately.
Problem management also matters if the issue recurs. Incident management restores service or processing; problem management investigates underlying causes and prevents recurrence. Repeated manual recovery without root-cause work can indicate that operational metrics are rewarding speed of restoration while ignoring chronic control failure.
An online service reports 99.98 percent availability. Traffic has grown rapidly, database utilization reaches 90 percent during peak periods, and one geographic region has little failover headroom. Management dismisses the concern because the service-level target is still being met.
Availability is backward-looking evidence; capacity management must also anticipate future demand. The auditor should examine capacity thresholds, forecasts, trend monitoring, stress testing, dependency limits, scaling behavior, and plans for peak events. If failover requires the surviving region to absorb full production load, normal utilization can conceal a resilience problem.
Service-level management should use meaningful measures. A monthly uptime percentage may hide brief outages during the highest-value transaction window or degraded performance that prevents users from completing work. Understand how the metric is calculated, what exclusions apply, whether data is complete, and whether the target reflects business impact.
The strongest CISA reasoning connects capacity to resilience. Redundancy that cannot carry required load after failure is not effective redundancy.
A finance team builds a critical forecasting process in a low-code platform and several spreadsheets because the central IT backlog is long. The process works well, but access is granted informally, formulas are changed without review, key files are stored in personal workspaces, and only one analyst understands the workflow.
Calling the solution “unauthorized” does not solve the risk. The auditor should evaluate business criticality, data sensitivity, access, change control, versioning, backup, documentation, continuity, input validation, output review, and ownership. The response should be proportionate. Low-risk end-user computing may need lightweight controls; a process feeding financial or regulatory decisions may need much stronger governance or migration to a managed platform.
This scenario tests whether you can distinguish the business need from the control gap. Simply prohibiting the tool can drive the process further underground. Sustainable remediation provides an approved path that preserves necessary agility while adding accountability and evidence.
A service desk meets its incident-resolution target by restoring failed services rapidly. The same storage, network, and application issues recur every month. Problem records are rarely opened, and change requests that would address root causes are deferred because they might affect availability metrics.
The auditor should examine whether incident, problem, and change management work as a system. Incident management minimizes business disruption. Problem management analyzes recurring or significant causes. Change management implements controlled remediation. Metrics that optimize only one function can produce perverse outcomes: fast closure with chronic instability.
Look for trend analysis, known-error records, root-cause methods, ownership, prioritization based on business impact, linkage between incidents and problems, approved corrective changes, and verification that remediation worked. If the organization cannot connect repeated incidents to durable fixes, management may be measuring activity rather than control effectiveness.
A logistics company backs up databases every night and stores copies in a separate account. The business impact analysis requires no more than 15 minutes of data loss and four hours of service downtime for the dispatch platform. Management claims the backup control is effective because every nightly job is green.
The evidence contradicts the recovery point requirement. A nightly backup alone cannot satisfy a 15-minute RPO unless another replication or capture mechanism reduces potential data loss. Even if RPO is addressed, RTO requires restoration of applications, configuration, secrets, network, identity, dependencies, and data within the required time.
A restoration test should measure actual outcomes, not merely whether a file can be read. Did the recovered environment start? Was data consistent? Were encryption keys available? Did interfaces reconnect? Could users authenticate? How long did each stage take? What manual steps or unavailable specialists were required?
The CISA lesson is to trace from business impact to technical evidence. Backup frequency is a design choice. RPO and RTO are business requirements. Testing proves whether the design can meet them.
A company conducts an annual disaster-recovery exercise. Staff receive the scenario two weeks in advance, the primary identity provider remains available, network teams preconfigure routes, and the exercise excludes third-party integrations. The final report says the recovery plan passed because the main application started in the secondary site.
Exercises should test realistic dependencies and decision making. Planned tests are useful, especially for safety, but over-scripted execution can provide false confidence. The auditor should assess scenario relevance, scope, participants, objectives, timing, recovery measurements, communication, vendor involvement, exceptions, issues, remediation, and retesting.
Dependencies deserve explicit mapping. DNS, identity, keys, certificates, network connectivity, cloud control planes, data replication, staff availability, facilities, vendors, telecommunications, and communications channels can all determine recovery success. A resilient application architecture is not a resilient business process if a single external dependency prevents operation.
Maturity grows through varied exercises: tabletop discussions, technical recovery tests, component failovers, communications exercises, and broader simulations. The appropriate mix depends on criticality and risk. The auditor should seek evidence that tests reveal weaknesses and that weaknesses lead to completed corrective action.
A company retains application, operating system, cloud, identity, database, and network logs. During an investigation, timestamps differ, some sources retain only seven days, privileged actions lack user attribution, and central monitoring receives only selected event types.
Logging is not an inventory exercise. Operational and security evidence needs completeness, consistent time, protected integrity, retention aligned with business and legal needs, useful identifiers, and reliable collection. The auditor should identify which events support critical control objectives and verify that those events can be correlated.
Centralization can improve analysis, but source configuration remains important. If an application never records a sensitive authorization decision, forwarding infrastructure cannot recreate it. Similarly, storing every event indefinitely can increase cost and privacy risk without improving assurance. Logging design should follow risk and investigative needs.
For CISA, ask what management must be able to prove. Can it identify who changed a critical configuration, when, under which approval, and with what result? Can it reconstruct a failed transaction? Can it demonstrate that a privileged action was reviewed? Evidence requirements make log design concrete.
A payment service runs across redundant infrastructure, but only two engineers know how to activate the alternate environment. One is on leave during an outage, and the other cannot access the emergency account because the recovery credentials were stored in a system affected by the same incident.
Business resilience includes people, access, communications, procedures, vendors, and decision authority. Technical redundancy fails if human and administrative dependencies are ignored. The auditor should review succession coverage, cross-training, emergency access, contact information, procedure currency, alternate communications, role assignments, and exercise participation.
This does not mean every employee must know every recovery task. Critical activities need sufficient depth so the loss of one person does not make the plan unusable. Exercises can reveal where knowledge is tacit rather than documented.
The scenario also demonstrates why recovery credentials and documentation should not depend solely on the production systems they are intended to recover. Secure alternate access must balance availability with the risk of creating an uncontrolled backdoor.
Use a six-step sequence. First, identify the business objective: reliable processing, accurate data, controlled change, recoverability, regulatory compliance, or another measurable outcome. Second, identify the lifecycle point at which the control should exist. Third, identify the responsible owner. Fourth, identify the evidence that would prove design and operation. Fifth, evaluate whether the observed condition threatens the objective. Sixth, select the auditor’s appropriate next action.
This sequence prevents technology bias. A scenario may mention a popular cloud, database, pipeline, or recovery product, but the exam usually tests a more durable principle: control requirements should be defined early, production changes should be authorized and traceable, migration should preserve data integrity, incidents should feed problem management, capacity should support continuity, and recovery capabilities should be tested against business objectives.
When two answers seem correct, compare scope and timing. One option may fix a symptom while another addresses the control objective. One may prescribe a tool before the auditor has verified the condition. One may belong to a later step. One may create an independence problem by making the auditor responsible for the control. State those distinctions explicitly during review instead of relying on intuition.
For a project, select one requirement and trace it from business case to approved requirement, design, implementation, testing, deployment, monitoring, and post-implementation review. Any broken link becomes a diagnostic question. If logging was required, where was the requirement approved? How was it designed? What test proved the right events were captured? Who monitors them after go-live? What evidence shows exceptions are acted upon?
For operations, choose one critical service and trace an outage. Which monitor detected it? Who was alerted? How was severity determined? What workaround restored service? Was data integrity validated? Was a problem record needed? Did a corrective change follow? Were capacity or configuration baselines updated? Did the event affect a resilience assumption?
For recovery, start from the BIA rather than the backup console. Write down maximum tolerable disruption, RTO, RPO, critical dependencies, minimum staffing, alternate processing needs, and communication obligations. Then map actual technical and procedural controls. If you cannot show evidence that each requirement has been tested, mark the gap.
These drills are valuable because they turn broad domains into chains of proof. CISA questions become easier when you habitually ask what evidence supports the conclusion and whether the control is positioned at the right point in the lifecycle.
Systems acquisition and implementation create promises: the solution will deliver certain benefits, protect certain information, satisfy certain control requirements, and be supportable. Operations and resilience determine whether those promises survive daily use, change, failure, growth, and disruption.
A strong CISA candidate can evaluate both sides of that handoff. You should recognize when a business case is incomplete, when agile evidence is valid even without traditional documents, when vendor responsibility is misunderstood, when testing misses operational readiness, when migration reconciliation is superficial, when emergency change bypass becomes normal, and when a post-implementation review never validates benefits. You should also recognize when batch success hides interface failure, when availability hides capacity risk, when shadow IT needs proportionate governance, when incident metrics discourage root-cause work, when backups do not meet recovery objectives, and when a disaster-recovery exercise proves only an idealized scenario.
The unifying habit is to follow the control from objective to evidence. Ask what management expected, what could prevent it, what control should exist, what proof is available, and whether the proof is strong enough to support the conclusion. That is the bridge between project assurance and operational resilience, and it is the reasoning these CISA domains are designed to test.
A business signs a managed-service agreement requiring 99.9 percent availability and four-hour incident notification. The contract does not define the measurement window, exclusions, source of availability data, severity model, or what constitutes notification. Months later, customer and provider dashboards report different uptime figures, and each side claims the other is calculating incorrectly.
The issue began during acquisition. A service-level commitment should be measurable and tied to evidence that both parties understand. The auditor should examine how requirements were translated into contract language, which systems produce the metrics, who validates them, how planned maintenance is handled, how disputes are resolved, and whether remedies or escalation are meaningful.
A percentage alone is not a complete control. If availability is measured from the provider’s internal load balancer while customers cannot authenticate through an external dependency, the metric may not represent the business service. If incident notification starts only after the provider internally classifies an event as severe, the customer may receive notice later than expected. Precise definitions make assurance possible.
This scenario also connects to operations. Once the contract is active, management should review service reports, investigate breaches, track recurring issues, and reassess whether the provider remains suitable. Acquisition defines the obligation; operations proves whether the obligation is being met.
An infrastructure group reports 97 percent patch compliance. The policy allows legacy systems with vendor restrictions to be marked “not applicable,” and those systems are omitted from the denominator. Several of them support critical business processes and have internet-reachable components.
The auditor should challenge population completeness and exception governance before celebrating the percentage. An exception may be legitimate, but risk does not disappear when a system cannot be patched. Management should identify compensating controls, owners, expiration or review dates, monitoring, network restrictions, upgrade plans, and accepted residual risk.
This example illustrates the relationship between configuration management, asset inventory, patch operations, and governance. If the inventory is incomplete, patch metrics are unreliable. If exception records are stale, management cannot tell whether compensating controls remain necessary or effective. If the systems are business critical, resilience planning must account for their increased failure and compromise risk.
A useful audit test is to work from independent asset sources toward the patch population rather than only from the patch tool outward. Cloud inventories, network discovery, procurement records, configuration repositories, endpoint management, and application ownership lists can reveal assets missing from the compliance report.
A regulated organization maintains a secondary recovery environment in another country. Technical tests show that systems can start there within the required RTO. Legal and privacy teams later determine that certain customer datasets cannot be processed in that jurisdiction under existing agreements.
This is a lifecycle integration failure. Resilience requirements should be reconciled with data classification, contractual obligations, regulatory constraints, and architecture before the recovery design is approved. A technically successful failover can still be unacceptable if it violates another business requirement.
The auditor should trace how recovery locations were selected, which datasets replicate there, what legal review occurred, what encryption and access controls apply, and whether alternate recovery strategies exist for restricted data. If management accepted the jurisdictional risk, verify that the acceptance was authorized and still valid.
The exam lesson is broader than data residency. Business continuity solutions must satisfy the same security, compliance, and governance obligations as normal operations unless an authorized exception explicitly changes them. Emergencies do not automatically suspend control objectives.
For both Domain 3 and Domain 4, classify controls by what they contribute. Preventive controls reduce the chance of an unwanted event: approval gates, least privilege, input validation, protected branches, capacity limits, or network restrictions. Detective controls reveal conditions or failures: reconciliations, monitoring, log review, drift detection, exception reports, and alerts. Corrective controls restore acceptable operation: rollback, incident response, problem remediation, restore procedures, and recovery plans.
A mature design usually uses layers. A release approval can prevent unauthorized deployment, a pipeline log can detect what actually occurred, and rollback can correct a failed release. Backup can enable correction after data loss, but it does not prevent the loss or detect corruption promptly. Monitoring can detect a capacity threshold, but it does not itself add capacity.
When a question presents multiple plausible controls, ask which control type addresses the stated gap. If the scenario says unauthorized changes are already occurring undetected, another preventive approval may be useful but detective evidence may be the immediate weakness. If the organization detects failures quickly but cannot restore service, the missing capability is corrective. This simple classification helps separate attractive but misaligned answers.
One of the most productive preparation exercises is to build a go-live handoff checklist from an auditor’s perspective. Include named service ownership, support model, monitoring, alert thresholds, incident and escalation paths, asset and configuration records, privileged-access ownership, backup and restore evidence, capacity baselines, vendor contacts, licensing, documentation, security exceptions, unresolved defects, data reconciliation, recovery procedures, service-level targets, and post-implementation review dates.
Then ask which items are mandatory for the system’s criticality and which can be accepted temporarily with documented risk. A low-impact internal tool and a life-safety system should not receive identical governance, but both need a deliberate decision. Risk-based control does not mean arbitrary control; it means the rigor is proportional and explainable.
During audit testing, compare handoff evidence with later operational reality. Did the named owner remain accountable? Were monitoring thresholds tuned after real traffic appeared? Were emergency changes incorporated into configuration baselines? Did recovery tests occur on schedule? Did unresolved defects become permanent exceptions? This longitudinal view reveals whether implementation controls actually established a sustainable operating environment.
When time is limited, reduce the scenario to five anchors: objective, lifecycle stage, owner, evidence, and next action. If the objective is data integrity during migration, do not get distracted by an unrelated uptime control. If the stage is pre-implementation readiness, do not choose a post-incident measure unless the question asks for contingency. If the auditor lacks reliable evidence, be cautious about jumping to a final conclusion. If the proposed answer makes the auditor own the operational control, examine independence.
Also watch absolute language. Controls rarely work because “all changes must be forbidden” or “every incident requires forensic imaging.” Strong governance is specific to risk, responsibility, evidence, and business need. The correct answer is often the one that restores those relationships rather than the one that sounds most severe.
Finally, preserve the connection between Domain 3 and Domain 4. Every implementation decision creates an operational consequence, and every recurring operational failure can reveal a weakness in design, transition, or governance. The candidate who sees that lifecycle connection is better prepared to evaluate systems as they actually exist rather than as isolated exam topics.
Popular posts
Recent Posts
