All questions
Question 1
An organization's problem management process requires a root cause analysis (RCA) for all P1 (critical) incidents. After a major database outage, no RCA is conducted because 'the team was too busy with other work.' The most significant risk of this gap is:
- The organization may fail its next ISO 27001 certification audit.
- Staff involved in the incident may not receive performance reviews on time.
- The underlying cause of the outage remains unaddressed, increasing the likelihood of recurrence and future disruptions to critical systems and financial operations. (correct answer)
- The incident ticket will remain in 'open' status indefinitely in the ticketing system.
Explanation: Without RCA, the root cause of a critical outage is unknown and unresolved - the same failure mechanism can cause another outage. Answer C is correct. Certification impacts (A) are secondary. Performance reviews (B) are unrelated. Open tickets (D) are an administrative issue.
Question 2
An IT team resolves incidents by restarting servers whenever applications crash, without investigating why the crashes occur. Over six months, the same servers are restarted 47 times. This approach reflects:
- Effective incident management since service is restored quickly each time.
- Effective incident management but failed problem management - the root cause of the crashes has not been investigated or resolved. (correct answer)
- A technical limitation of the servers that cannot be resolved.
- An appropriate use of the change management process to address recurring issues.
Explanation: Quick restarts demonstrate responsive incident management. However, 47 restarts without root cause investigation is a clear problem management failure - the recurring crashes indicate an unresolved underlying issue. Answer B is correct. Quick restoration alone is not sufficient (A). The frequency suggests a resolvable problem (C). Restarts are not change management (D).
Question 3
During an audit, an IT auditor reviews the incident log and finds that several high-severity incidents affecting the financial reporting system were not logged in the incident management system. The primary risk of unlogged incidents is:
- The organization loses visibility into system reliability patterns, management cannot make informed decisions, and root cause analysis cannot be performed on untracked issues. (correct answer)
- The IT team will exceed its incident handling capacity.
- The organization will fail its SOC 2 audit automatically.
- Employees who experienced the incidents may file complaints about poor service.
Explanation: Unlogged incidents create blind spots - management cannot see patterns, cannot perform trend analysis, and cannot trigger problem management for recurring issues. Answer A is correct. Capacity (B), automatic SOC failures (C), and employee complaints (D) are not the primary risks.
Question 4
An organization's incident management process requires that all resolved incidents be reviewed within 5 business days to confirm the resolution is effective and the incident has not recurred. This post-resolution review primarily supports which objective?
- Verifying that the fix was effective and triggering problem management investigation if the incident recurs within the review window. (correct answer)
- Ensuring that IT staff properly documented the incident resolution for billing purposes.
- Confirming that the affected user has submitted a satisfaction rating for the support experience.
- Archiving the incident record for regulatory compliance purposes.
Explanation: Post-resolution review confirms fix effectiveness and catches early recurrences that should trigger problem management - connecting incident and problem management processes. Answer A is correct. Billing documentation (B), satisfaction ratings (C), and archiving (D) are administrative activities that are not the primary purpose.
Question 5
An organization's ITSM platform automatically creates a problem record when three or more incidents with the same category and affected system are logged within a 30-day period. This automation is designed to:
- Reduce the number of incident tickets by merging related incidents.
- Proactively trigger problem management investigation when patterns suggest an underlying systemic issue, preventing continued reactive incident handling. (correct answer)
- Alert end users that their issues are part of a known pattern.
- Calculate the financial cost of systemic problems for management reporting.
Explanation: Automated problem record creation based on incident patterns is a proactive problem management trigger - identifying systemic issues before they cause further damage, rather than waiting for manual escalation. Answer B is correct. Merging tickets (A) is a different function. User alerts (C) and cost calculation (D) are secondary functions.
Question 6
During an audit, an organization claims its incident management process is effective because 'issues get fixed.' The auditor should evaluate this claim by reviewing:
- Employee satisfaction survey results about IT support quality.
- The IT department's headcount and training certifications.
- The organization's IT strategic plan for planned system improvements.
- Incident logs, SLA compliance rates, recurring incident patterns, root cause analysis completion rates, and evidence of permanent fixes implemented. (correct answer)
Explanation: Auditing incident management effectiveness requires objective evidence: ticket data, SLA compliance, recurrence patterns, RCA completion, and fix implementation - not anecdotal claims. Answer D is correct. Satisfaction surveys (A), headcount (B), and strategic plans (C) do not directly measure incident management process effectiveness.
Question 7
An auditor evaluating incident management controls for a financial services company finds that the company has no documented incident response procedures. The primary risk is:
- Regulators will impose fines for the absence of documentation.
- IT staff will be unable to use the incident management system.
- The company will be unable to obtain cyber insurance.
- During an actual incident, the response will be ad hoc and inconsistent, potentially increasing resolution time, escalating impact, and missing critical response steps. (correct answer)
Explanation: Without documented procedures, incident response depends on individual knowledge and improvisation - leading to inconsistent, slower, and incomplete responses that increase damage. Answer D is correct. Regulatory fines (A) may follow but are secondary. System usability (B) is unrelated. Insurance (C) is a separate consideration.
Question 8
An organization's incident management SLA requires that P1 incidents be resolved within 4 hours. The auditor reviews the incident log and finds that 35% of P1 incidents exceeded the 4-hour SLA during the past year. The auditor should:
- Accept this as reasonable since some incidents are more complex than others.
- Recommend that the SLA be extended to 8 hours to improve compliance.
- Flag this as a control deficiency - a 35% SLA breach rate for critical incidents indicates systemic weaknesses in the incident response capability that require investigation and remediation. (correct answer)
- Accept this if management provides an explanation for each exceeded SLA.
Explanation: A 35% SLA breach rate for the highest-priority incidents is a significant finding that indicates systemic problems with incident response capacity, prioritization, or process - not isolated exceptions. Answer C is correct. Accepting high breach rates (A) ignores the control failure. Extending the SLA (B) masks the problem. Individual explanations (D) do not address the systemic issue.
Question 9
During an audit of incident management controls, an auditor finds that critical system incidents are not escalated to senior management until they have been unresolved for more than 48 hours. The primary risk of this escalation policy is:
- Senior management will receive too many notifications about minor incidents.
- Critical incidents may remain unresolved for up to 48 hours without senior management awareness, delaying resource allocation and potentially exceeding recovery time objectives. (correct answer)
- The IT team will spend too much time preparing escalation reports for management.
- The escalation policy may conflict with the organization's SLA commitments to vendors.
Explanation: A 48-hour escalation delay for critical incidents means management is unaware and cannot allocate additional resources for up to two days - potentially well beyond the RTO for critical systems. Answer B is correct. Over-notification (A) is not the risk for critical incidents. Report preparation (C) is a minor concern. Vendor SLAs (D) are a separate consideration.
Question 10
A company experiences a system outage and the IT team restores service within the SLA timeframe. However, the same outage occurs again three weeks later. This pattern most likely indicates a failure in:
- Problem management - the root cause was not identified and eliminated after the first incident, allowing the underlying issue to cause recurrence. (correct answer)
- Incident management - the first restoration took too long.
- Change management - the outage was caused by an unapproved change.
- Backup and recovery - the system was not properly backed up before the outage.
Explanation: Recurring incidents indicate that incident management restored service but problem management failed - the root cause was not found and fixed, allowing the same underlying issue to cause another outage. Answer A is correct. The first restoration met SLA (B). There is no evidence of a change (C) or backup issue (D).
Question 11
Which of the following represents an effective integration between incident management and change management processes?
- When problem management identifies a root cause requiring a fix, the fix is implemented through the formal change management process to ensure it is authorized, tested, and documented. (correct answer)
- Change management should approve all incident resolutions before service is restored.
- Incident management and change management should operate independently to avoid delays.
- Only problem managers are authorized to initiate change requests.
Explanation: The formal link between problem management and change management ensures that fixes identified through root cause analysis are implemented in a controlled, authorized manner - preventing rushed fixes that could cause new problems. Answer A is correct. Pre-approval of all incident restorations (B) would cause unacceptable delays. Independence (C) creates gaps. Change initiation is not limited to problem managers (D).
Question 12
Which of the following is the most important information to capture in an incident record to support effective problem management?
- The name of the end user who first reported the incident.
- Symptoms, affected systems, timeline of events, steps taken to resolve, resolution method, and root cause if identified. (correct answer)
- The cost of IT staff time spent on the incident.
- The number of users affected by the incident.
Explanation: Comprehensive incident records with symptoms, timelines, and resolution details provide the foundation for problem management root cause analysis - enabling pattern recognition and systematic investigation. Answer B is correct. Reporter name (A), cost (C), and user count (D) are supplementary data that do not support root cause investigation.
Question 13
Which of the following is a key control that helps ensure incidents are escalated appropriately when they cannot be resolved within defined timeframes?
- Documented escalation paths and timeframes that automatically trigger notification of senior staff and management when incidents breach defined resolution windows. (correct answer)
- Requiring all IT staff to carry mobile phones so they can be reached at any time.
- Posting the IT team's organizational chart in the server room.
- Conducting monthly incident management training for IT staff.
Explanation: Documented escalation paths with defined triggers ensure that unresolved incidents automatically escalate to higher levels of authority, ensuring resources and management attention are applied before incidents cause unacceptable disruption. Answer A is correct. Mobile phones (B), org charts (C), and training (D) are supporting elements but not the escalation control itself.
Question 14
Which of the following metrics is most useful for evaluating the effectiveness of an incident management process?
- Mean time to resolve (MTTR) - the average time from incident detection to service restoration. (correct answer)
- Mean time between failures (MTBF) - the average time between system failures.
- Total number of incidents logged per month.
- Percentage of incidents reported by end users versus automated monitoring.
Explanation: MTTR directly measures incident management effectiveness - how quickly the team restores service after an incident. Answer A is correct. MTBF (B) measures reliability, not incident management effectiveness. Total incidents logged (C) measures volume, not resolution effectiveness. Reporting source (D) is a detection metric.
Question 15
In IT service management (ITSM), what is the primary distinction between 'incident management' and 'problem management'?
- Incident management is performed by senior staff; problem management is performed by junior staff.
- Incident management addresses security breaches; problem management addresses software defects.
- Incident management is a reactive process; problem management is a proactive process only.
- Incident management focuses on restoring normal service as quickly as possible; problem management focuses on identifying and eliminating the root cause to prevent recurrence. (correct answer)
Explanation: Incident management prioritizes speed of restoration - getting users back to work. Problem management digs deeper to find and eliminate the underlying root cause so the incident does not recur. Answer D is correct. Seniority (A) and domain scope (B) are not the distinguishing factors. Problem management can be both reactive and proactive (C).
Question 16
A company experiences a ransomware attack that encrypts 60% of its production data. The incident response team's first priority should be:
- Immediately paying the ransom to restore data as quickly as possible.
- Notifying all customers that their data may be compromised.
- Containing the attack by isolating affected systems to prevent further spread, then beginning the investigation and recovery process. (correct answer)
- Conducting a root cause analysis to determine how the ransomware entered the network.
Explanation: The first incident response priority is containment - isolating affected systems to stop the ransomware from spreading further. Investigation, notification, and recovery follow containment. Answer C is correct. Paying ransom (A) is a last resort that doesn't guarantee recovery. Customer notification (B) comes after containment and assessment. Root cause analysis (D) follows containment.
Question 17
A 'known error' in ITSM problem management refers to:
- A documented security vulnerability that has been publicly disclosed.
- An error in the incident ticketing system that causes tickets to be misrouted.
- An error in the organization's financial statements identified by the external auditors.
- A problem that has been diagnosed with an identified root cause and a known workaround, but for which a permanent fix has not yet been implemented. (correct answer)
Explanation: A known error is a formally documented problem state where the root cause and a workaround are identified - it is tracked until a permanent fix (change) is implemented. Answer D is correct. Security vulnerabilities (A), ticketing errors (B), and financial statement errors (C) are not the ITSM definition.
Question 18
Which of the following best describes a 'workaround' in ITSM incident and problem management?
- A permanent fix that eliminates the underlying cause of repeated incidents.
- A temporary patch applied by the IT team to a software vulnerability.
- A manual process that replaces automated system functionality indefinitely.
- A temporary solution that reduces the impact of an incident or problem until a permanent fix can be implemented. (correct answer)
Explanation: A workaround is a temporary measure - it mitigates impact but does not fix the underlying cause. It buys time until a proper solution is developed and implemented. Answer D is correct. A permanent fix (A) eliminates the need for a workaround. A security patch (B) may be a fix, not a workaround. Manual processes (C) may be workarounds but the definition is broader.
Question 19
An organization's problem management process uses trend analysis of incident data. The primary purpose of this analysis is to:
- Determine whether the IT team is meeting its staffing targets.
- Identify recurring patterns or categories of incidents that may indicate underlying systemic problems requiring root cause investigation. (correct answer)
- Calculate the total cost of IT outages for financial reporting purposes.
- Provide data for the annual IT performance review.
Explanation: Trend analysis of incident data reveals patterns - the same system failing repeatedly, the same type of error occurring frequently - that signal underlying problems requiring problem management attention. Answer B is correct. Staffing (A), cost calculation (C), and performance reviews (D) are secondary uses.
Question 20
After a major security incident, an organization conducts a post-incident review. The primary purpose of this review is to:
- Determine which employees are responsible for the incident for disciplinary action.
- Satisfy the insurance company's requirements for incident reporting.
- Document the incident for inclusion in the annual IT report to the board.
- Understand what happened, why it happened, what was done well, what could be improved, and what actions will prevent recurrence. (correct answer)
Explanation: A post-incident review (also called a post-mortem or lessons learned) is focused on understanding and improvement - not blame - covering the full incident timeline, response effectiveness, and preventive actions. Answer D is correct. Blame assignment (A) is counterproductive. Insurance reporting (B) is a compliance activity. Board reporting (C) may follow but is not the review's primary purpose.