Incident Management

Worth 30% of the ISACA CISM (CISM) exam. CertClue has 461 questions on this objective.

What this objective covers

Events, Incidents and Breaches: Core Concepts

Precision in this vocabulary decides several exam questions. An event is any observable occurrence; most events are routine. A security event is one with security relevance, such as a failed logon or a blocked connection. An incident is an event or series of events that actually threatens or harms confidentiality, integrity or availability, or violates policy, and it is the point at which the response process is invoked. A breach is a specific outcome in which protected information is confirmed to have been accessed, disclosed or taken by an unauthorised party, and it usually carries notification duties under the applicable requirements. The distinction matters operationally: incidents are declared on suspicion so that response can begin, whereas a breach is a determination made on evidence, and calling something a breach before the evidence supports it creates obligations and reputational damage the facts may not justify.

Exam tip. Declare an incident on reasonable suspicion so response can start, but declare a breach only when the evidence supports it. Confirming disclosure prematurely creates obligations the facts may not justify.

The Incident Response Plan and Team Structure

Preparation is the phase that determines whether everything after it works. The incident response plan defines scope, roles and authority, classification criteria, escalation paths, communication requirements, evidence handling rules and the criteria for invoking continuity plans. It must state who has authority to make disruptive decisions such as isolating a production system or taking a service offline, because arguing about that authority in the middle of an incident costs the time that matters most. The team is cross-functional by design: security and IT provide technical handling, but legal, communications, human resources, business owners and executive sponsors all have defined roles. External support such as forensic specialists and legal counsel should be contracted in advance, since procurement during an incident is slow. Contact lists, out-of-band communication channels and offline copies of the plan matter because the incident may have taken down the systems you would normally use.

Exam tip. When responders hesitate or escalate slowly, the root cause is usually undefined authority or unrehearsed procedures. Preparation, not detection technology, is the corrective action.

Incident Classification and Severity Assignment

Classification puts an incident into a category, such as malicious code, unauthorised access, data disclosure, service disruption or misuse, so that the right playbook and the right specialists are engaged. Severity is a separate judgement about how bad this instance is, driven by business impact: which processes are affected, how much data and of what sensitivity, how many customers, whether the impact is spreading, and what the obligation and reputational consequences look like. Severity should be assigned against pre-agreed criteria rather than negotiated during the incident, because consistent criteria are what make escalation automatic and response times defensible. Severity is provisional and must be revisited as facts emerge, in both directions. Under-classification delays the resources an incident needs, while reflexive over-classification exhausts senior attention and teaches people to ignore the highest level when it is real.

Exam tip. Severity is set by business impact, not by technical sophistication. A simple failure that halts a critical revenue process outranks an advanced attack on an isolated test system.

Detection, Triage and Reporting Channels

Detection determines how long an attacker operates unobserved, and dwell time is the variable most closely linked to eventual impact. Sources are both technical and human: monitoring and correlation across logs, endpoint and network telemetry, integrity checking, and equally staff, customers and third parties who notice something wrong. Human reporting only works if reporting is easy, well known, available around the clock and free from fear of blame, which is why punitive handling of honest mistakes materially reduces detection capability. Triage confirms whether a signal is a genuine incident, assigns initial classification and severity, and starts the record. A single well-publicised point of contact prevents reports scattering across individual inboxes where they die. From the first moment, every action, decision, time and finding is logged, because that record supports later analysis, evidence handling and reporting to governance.

Exam tip. If staff notice problems but do not report them, fix the reporting channel and the culture around it. Adding more monitoring technology does not solve a human reporting failure.

Escalation, Declaration and Decision Authority

Escalation is the mechanism that puts an incident in front of the people who can authorise disruptive action, commit resources and speak for the organisation. It should be criteria-driven: defined severity levels, categories or trigger conditions automatically escalate, so no responder has to decide alone whether an executive would want to be woken. Escalation runs on two tracks. Functional escalation brings in more or different technical capability; hierarchical escalation brings in more authority. Confusing them wastes time, because a responder who needs permission to take a revenue system offline does not need another analyst. Declaration formally starts the response and may also start continuity invocation. Certain triggers should always escalate regardless of apparent size: suspected disclosure of sensitive information, suspected insider involvement, safety implications, suspected compromise of the response tooling itself, and anything likely to attract external attention.

Exam tip. When a responder needs permission rather than help, that is hierarchical escalation. Expect scenarios where the delay was caused by escalating to more technical staff instead of to someone with authority.

Incident Communication and Stakeholder Management

Communication during an incident is a controlled activity with a single accountable owner, because inconsistent or premature statements do lasting damage that the technical response cannot undo. Internally, staff need to know what to do and what not to say; executives need impact, actions and decisions required, not packet-level detail. Externally, customers, partners, insurers and the parties identified in the applicable requirements each have distinct needs and timeframes, and one designated spokesperson should speak publicly so the message stays consistent. Communicate what is known, what is not yet known and what happens next, and avoid speculation about cause or scope while the facts are still moving, because a retraction is more damaging than an initial statement that was appropriately cautious. Where information may have been disclosed, engage legal and privacy advisers early so notification obligations are assessed against the applicable requirements from the start rather than discovered late.

Exam tip. Where information may have been disclosed, involve legal and privacy advisers early and communicate under the applicable requirements. Never let technical staff make public statements about cause or scope.

Containment Strategies and Their Trade-offs

Containment stops the incident spreading and limits damage, and it is the first priority once an incident is confirmed, subject to safety. It splits into short-term containment, which acts immediately to stop the bleeding through isolation, account disabling, blocking or traffic filtering, and long-term containment, which applies more durable measures that let the business keep operating while eradication is prepared. Every containment decision is a trade-off the security manager must frame for the business: aggressive isolation stops the spread but may halt revenue-generating services and may destroy volatile evidence; a quieter approach preserves evidence and continuity but allows the attacker more time. Powering off a compromised system loses memory-resident evidence, which is why isolation from the network is usually preferred to shutdown. Where the attacker may be watching, avoid alerting them until a coordinated containment can be executed across all footholds at once.

Exam tip. Once an incident is confirmed, contain before eradicating and isolate rather than powering off, unless safety requires otherwise. Business-impacting containment needs the authority named in the plan.

Eradication, Recovery and Return to Normal Operations

Eradication removes the cause and the attacker's means of return: malicious artefacts, unauthorised accounts and access, altered configurations and the underlying weakness that allowed entry. Recovery restores systems and data to normal operation and validates that they are clean and functioning correctly before they carry live business again. Restoring from backup only works if the restore point predates the compromise and the backup itself is verified, which is why establishing the timeline during analysis is a prerequisite for choosing it. Sequence restoration by business priority using the recovery objectives agreed in continuity planning, not by whichever system is easiest to bring back. After restoration, monitor the recovered environment at heightened sensitivity for a defined period, because the most common failure at this stage is declaring closure while a foothold remains and the incident recurs a fortnight later. Closure is a decision with criteria, taken by the incident owner.

Exam tip. Recovery is not complete when the service is back. It is complete when the system is validated clean, monitored, and the exploited weakness has been closed so the same route cannot be reused.

Evidence Handling, Forensics and Chain of Custody

Any incident may later support disciplinary action, an insurance claim, a contractual dispute or a legal proceeding, so evidence is handled properly from the start because you cannot retrospectively make it admissible. Collect in order of volatility, capturing memory and running state before disk and archived data, since the most fragile evidence disappears first. Work on forensic copies rather than originals, and use verification values so that the copy can be shown to match the source and to be unaltered since. Chain of custody records who collected each item, when, from where, and every subsequent transfer and access, because a gap in that record undermines the evidence regardless of its content. Only competent, authorised people should collect and analyse, and where in-house capability is thin, engaging specialists early beats damaging the evidence first. Retention of evidence continues until legal, privacy and business advisers agree it can be released.

Exam tip. Preserve evidence before making changes, capture volatile data first, and keep an unbroken chain of custody. If in doubt about competence, engage specialists rather than improvising.

Post-Incident Review, Root Cause and Corrective Action

The post-incident review is where the organisation converts an expensive event into improvement, and skipping it guarantees repetition. Hold it soon enough that memory is fresh, with everyone involved, and run it without blame so that people describe what actually happened rather than what makes them look competent. The review covers the timeline, what was detected and when, what worked, what did not, and what the root cause was, distinguishing the technical trigger from the underlying process or governance failure that allowed it. Draw the distinction between correction, which fixes this instance, and corrective action, which addresses the cause so the class of incident does not recur. Every corrective action needs an owner, a date and follow-up verification, and it should be tracked in the risk register or programme plan rather than in the incident report where it will be forgotten. Feed the findings into the risk assessment, the awareness programme and the response plan itself.

Exam tip. Rebuilding the affected system is a correction, not a corrective action. Expect the correct answer to address the underlying process or governance failure that allowed the incident.

Business Continuity Planning and the Business Impact Analysis

Business continuity planning keeps critical business processes running through a disruption, whatever the cause. It starts with a business impact analysis, which identifies critical processes, their dependencies on people, systems, suppliers and facilities, and how impact grows over time. The BIA is what produces the recovery objectives: maximum tolerable downtime, the point beyond which the business cannot survive the outage; recovery time objective, the target for restoring the process, which must sit inside the maximum tolerable downtime; and recovery point objective, the amount of data loss the business can accept, which drives backup and replication frequency. These are business decisions made by process owners on the basis of impact, not technical targets chosen by IT. Plans then cover alternate ways of working, alternate facilities, staff arrangements, supplier dependencies and the criteria and authority for invocation.

Exam tip. Recovery objectives come from the business impact analysis and are owned by business process owners. Any option where IT sets RTO or RPO on technical grounds is wrong.

Disaster Recovery, Recovery Sites and Plan Testing

Disaster recovery is the technology-focused subset of continuity: restoring the systems, data and infrastructure that the business processes depend on, to meet the objectives continuity planning set. The recovery strategy follows from those objectives, and site choice is the classic trade-off between cost and speed, from a hot site that can take over almost immediately, through warm and cold sites, to reciprocal or cloud-based arrangements. Backups sit underneath everything, and the exam consistently rewards the answer that a backup is not a recovery capability until a restore has been tested successfully. Plans decay as systems, staff and suppliers change, so they are tested on a schedule and after significant change, escalating from document review and walkthrough to tabletop, simulation, parallel testing and finally full interruption, which is the most conclusive and the most disruptive. Every test produces findings that update the plan.

Exam tip. Full interruption testing gives the strongest assurance and carries the greatest risk, so it needs executive approval. Tabletop exercises are the usual right answer when disruption must be avoided.

Practice questions

Free, with the answer and the reasoning. No account needed.

1. A newly appointed information security manager discovers that the incident response plan was written by a consultant three years ago and has never been revised. What should the manager do FIRST?

  • A. Purchase an incident response platform to automate the workflow
  • B. Schedule incident response training for the service desk
  • C. Distribute the existing plan to all staff so everyone has a copy
  • D. Review the plan against the current business and technology environment to identify where it no longer appliescorrect

A plan that no longer reflects how the organization actually operates will fail at the moment it is needed, so the manager must first establish the gap between the documented plan and current reality before doing anything with it. Distributing the plan is the tempting choice because it feels like fast progress, but circulating instructions that reference systems and providers that no longer exist actively increases the chance of a wrong action during an incident.

2. An information security manager is preparing to conduct a business impact analysis across twelve departments. Which approach will produce the MOST reliable results?

  • A. Distributing a questionnaire and consolidating the returned responses
  • B. Structured interviews with process owners using consistent criteria, with impacts quantified over defined time intervalscorrect
  • C. Asking the technology team to estimate impacts for each department's systems
  • D. Using impact figures published for comparable organizations in the sector

Consistent criteria applied through structured interviews produce comparable, defensible impact figures that can be ranked across departments. A questionnaire is the efficient looking alternative, but unmoderated responses vary wildly in interpretation and typically return everything marked critical, which cannot be used to prioritize anything.

3. What should PRIMARILY determine how often a specific continuity plan is exercised?

  • A. The availability of the recovery vendor's test windows
  • B. The frequency required by the previous external audit
  • C. The number of staff available to participate
  • D. The criticality of the processes it covers and the rate of change affecting themcorrect

Testing frequency should track how much damage failure would cause and how quickly the environment moves away from what the plan assumes. Audit frequency is the tempting anchor because it is concrete and externally imposed, but it sets a floor for compliance rather than a schedule matched to the organization's actual risk.

Work the whole objective

The full ISACA CISM bank, the study notes behind these summaries, and a readiness score that tells you which objective to revise next. Free, no paid tier.

Take the free ISACA CISM practice test

The other ISACA CISM objectives