CPA (ISC) • INFORMATION SYSTEMS

Evaluate Incident And Problem Management

How organizations systematically detect, resolve, and learn from IT disruptions to protect financial reporting integrity.

Historical Context & Motivation

The discipline of incident and problem management did not spring from information technology alone; it evolved at the intersection of enterprise risk management, audit standards, and the growing dependence of financial reporting on complex IT environments. In the 1970s and 1980s, mainframe-era organizations handled system failures in an ad hoc fashion—technicians fixed things as they broke, with limited documentation. As the Sarbanes-Oxley Act of 2002 (SOX) imposed rigorous internal-control requirements on publicly traded companies, it became clear that uncontrolled IT disruptions could compromise financial statement reliability. CPA auditors, already evaluating internal controls over financial reporting (ICFR), had to extend their assessments into the information systems domain, making structured incident and problem management a cornerstone of IT general controls (ITGCs).

1989
ITIL v1 Published
The UK government publishes the first Information Technology Infrastructure Library (ITIL), formalizing incident management as a dedicated IT service management process and separating it from problem management for the first time.
2002
Sarbanes-Oxley Act Enacted
SOX Section 404 mandates management assessment and external audit of internal controls over financial reporting, pushing IT incident logging and resolution into the auditor's scope.
2004
COSO ERM Framework
The Committee of Sponsoring Organizations publishes its Enterprise Risk Management framework, linking IT operational events—including incidents—to organizational risk appetite and strategic objectives.
2012
COBIT 5 Released
ISACA releases COBIT 5, which includes the DSS02 (Manage Service Requests and Incidents) and DSS03 (Manage Problems) processes, providing CPA-friendly control objectives for evaluating IT service continuity.
2019
ITIL 4 & Modern Practices
ITIL 4 reconceptualizes incident and problem management within a service value system, emphasizing continual improvement and integration with DevOps, reflecting the pace at which modern financial systems operate.

The central question for CPA candidates is therefore not merely technical—how does an IT department fix a server outage?—but rather evaluative: Are the organization's processes for detecting, recording, escalating, resolving, and learning from IT disruptions designed and operating effectively enough to safeguard the integrity of financial data? This lesson equips you to answer that question with the analytical rigor expected on the CPA ISC examination.

Core Principles & Definitions

Before evaluating incident and problem management controls, a CPA must command precise definitions. The terms incident and problem are related but distinct. An incident is any unplanned interruption to an IT service or a reduction in the quality of an IT service—think of a payroll application crashing on payroll-processing day. A problem is the underlying, often unknown, root cause of one or more incidents—in our example, perhaps the payroll application crashed because a recent database patch introduced a memory leak. Incident management aims to restore normal service as quickly as possible, while problem management seeks to identify and eliminate root causes so that incidents do not recur.

1

Incident Management

The reactive process of detecting, logging, categorizing, prioritizing, and resolving unplanned service disruptions to restore normal operations within agreed service-level targets.
2

Problem Management

The proactive and reactive process of performing root cause analysis (RCA) to prevent recurrence of incidents and minimize the impact of incidents that cannot be prevented.
3

Known Error Database (KEDB)

A repository that stores details of identified problems and documented workarounds, enabling faster incident resolution and forming part of the organizational knowledge base.
4

Service Level Agreements (SLAs)

Formally documented commitments defining response and resolution timeframes for incidents of varying severity, against which the effectiveness of incident management is measured.
5

Escalation Procedures

Defined pathways—both functional (technical) and hierarchical (management)—that ensure incidents receive adequate expertise and executive attention when initial resolution attempts fail.
KEY TAKEAWAY
Think of incident management like the emergency room at a hospital: the immediate goal is to stabilize the patient (restore IT service) and record the symptoms. Problem management is the specialist follow-up—the cardiologist who runs diagnostic tests to find out why the patient collapsed and prescribes treatment to prevent it from happening again. A CPA evaluator must assess whether both the ER and the specialist clinic are adequately staffed, documented, and governed.

Visual Explanation — The Incident & Problem Management Lifecycle

The left column depicts the five stages of incident management (detect → categorize → investigate → resolve → close). The right column shows the five stages of problem management (identify → RCA → workaround → permanent fix → close). Note the cross-links: recurring incidents escalate into problem records (pink dashed arrow), and workarounds discovered during problem management feed back into incident resolution (green dashed arrow). A CPA evaluator assesses whether each stage is properly designed and documented.

As the diagram illustrates, incident and problem management are parallel but interconnected processes. The CPA evaluator's role is to assess each stage for the presence of adequate controls: Is every incident logged automatically or is there discretion that allows gaps? Are priority levels assigned using a documented matrix rather than subjective judgment? Are escalation thresholds defined and enforced? Does root cause analysis follow a structured methodology such as the Ishikawa (fishbone) diagram or the Five Whys technique? Each of these questions maps directly to a testable control objective.

How It Works — The Evaluation Framework

While incident and problem management are not primarily mathematical disciplines, the CPA evaluator relies on quantitative metrics and qualitative control assessments to determine whether these processes are operating effectively. Several key performance indicators (KPIs) underpin the evaluation, and understanding how they are computed allows the auditor to identify control deficiencies with precision.

INCIDENT RESOLUTION RATE
IRR = (Incidents Resolved Within SLA ÷ Total Incidents) × 100%
Where IRR measures the percentage of incidents closed within agreed service levels. An IRR consistently below 90% may signal inadequate staffing, unclear escalation procedures, or poorly defined SLA targets.
MEAN TIME TO RESOLVE (MTTR)
MTTR = Σ (Resolution Timeᵢ) ÷ n
Where Resolution Timeᵢ is the elapsed time from incident detection to service restoration for incident i, and n is the total number of incidents in the measurement period. Trending MTTR by severity level helps auditors detect deterioration in response capability.
RECURRENCE RATE
RR = (Incidents Linked to Known Problems ÷ Total Incidents) × 100%
A high recurrence rate indicates that problem management is failing to eliminate root causes—a significant control deficiency for any organization reliant on IT-dependent financial processes.

Beyond metrics, the CPA evaluator applies a controls-based evaluation framework that mirrors the general approach to testing IT general controls. The evaluator first identifies the relevant control objectives from a recognized framework (such as COBIT DSS02 and DSS03), then assesses control design by reviewing policies, procedures, and system configurations. Finally, the evaluator tests operating effectiveness by examining a sample of incident and problem records, verifying that documentation is complete, escalation protocols were followed, and resolutions were applied within SLA windows. Any gap between the stated policy and actual practice constitutes a control deficiency, which is then assessed for severity—ranging from a simple deficiency through a significant deficiency to a material weakness, depending on the likelihood and magnitude of financial misstatement that could result.

📝 CPA EXAM TIP
On the ISC examination, expect scenarios where you must distinguish between a deficiency in design (e.g., no escalation policy exists) and a deficiency in operating effectiveness (e.g., escalation policy exists but 40% of critical incidents were not escalated). Both require different remediation strategies.

Detailed Breakdown — Incident Severity & Priority Matrix

A well-designed incident management process classifies every incident along two dimensions: impact (the breadth and business significance of the disruption) and urgency (how quickly the business requires a resolution). The intersection of these two dimensions determines the priority level, which in turn governs SLA targets, resource allocation, and escalation paths. A CPA evaluating incident management must verify that this matrix exists, is consistently applied, and aligns with the organization's risk tolerance—particularly for incidents affecting financially significant applications such as the general ledger, accounts payable/receivable, and treasury management systems.

The priority matrix maps Impact (vertical axis) against Urgency (horizontal axis) to yield five priority levels from P1 (Critical) to P5 (Planning). Each cell includes the typical SLA resolution window. CPA evaluators verify that the organization's matrix is formally documented, that SLA targets are reasonable relative to business risk, and that incident records consistently reflect the assigned priority.
Representative incidents by priority level and their potential impact on financial reporting
PriorityExample ScenarioFinancial Reporting Impact
P1 – CriticalERP general ledger module down during month-end closeDirect: journal entries cannot be posted, closing delayed, potential for misstated period-end balances
P2 – HighAccounts payable batch processing fails for one supplierModerate: delayed payments, accrual misstatement, potential vendor relationship issues
P3 – MediumReporting dashboard intermittently slowIndirect: management review controls hampered, but underlying data integrity preserved
P4 – LowSingle user unable to access training environmentMinimal: no impact on production financial data or processing

Worked Example — Evaluating Incident Management at Zeta Corp

Consider Zeta Corp, a publicly traded manufacturing firm whose financial statements are subject to an integrated audit. You are a CPA evaluating IT general controls. During your review of Zeta Corp's incident management process for the fiscal year ended December 31, you obtain the following data from the IT service management (ITSM) system: 1,200 total incidents logged; 1,020 resolved within SLA targets; 180 that breached SLA; 60 P1-Critical incidents, of which 48 were resolved within the one-hour SLA window; and 15 incidents that recurred three or more times and were eventually linked to a single known problem (a recurring database timeout). Your task is to evaluate these metrics and determine whether any control deficiencies exist.

Evaluating Zeta Corp's Incident & Problem Management Controls
1
Step 1 — Compute the Overall Incident Resolution Rate (IRR)Apply the formula: IRR = (Incidents Resolved Within SLA ÷ Total Incidents) × 100%. With Zeta Corp's data: IRR = (1,020 ÷ 1,200) × 100% = 85.0%. This falls below the commonly accepted benchmark of 90%, signaling a potential weakness in the overall incident resolution process.
IRR = 85.0% — below the 90% benchmark
2
Step 2 — Analyze P1-Critical Resolution RateFor the most severe incidents: P1 IRR = (48 ÷ 60) × 100% = 80.0%. Twelve critical incidents breached the one-hour SLA, meaning financially significant systems (likely the ERP or trading platforms) were impaired for extended periods. This warrants deeper investigation into escalation compliance.
P1 IRR = 80.0% — 12 critical SLA breaches
3
Step 3 — Assess the Recurrence RateThe 15 recurring incidents linked to the database timeout problem suggest that problem management did not resolve the root cause promptly. Computing the recurrence contribution: (15 ÷ 1,200) × 100% = 1.25%. While the percentage appears small, the qualitative significance is high because these 15 incidents likely affected a financially critical system repeatedly—exactly the scenario where a CPA should probe deeper into the Known Error Database and whether a Request for Change (RFC) was submitted.
Recurrence Rate = 1.25% — qualitatively significant due to financial system impact
4
Step 4 — Identify Control DeficienciesBased on the analysis, two potential deficiencies emerge. First, the overall and P1 IRRs suggest a deficiency in operating effectiveness of the incident resolution control—the policy exists (design is adequate), but actual performance falls short. Second, the recurring database timeout incidents that were not eliminated through problem management suggest either a design deficiency (no formal trigger for escalating recurring incidents to problem management) or an operating effectiveness deficiency (the trigger exists but was not followed). The auditor should examine Zeta Corp's problem management policy to make this determination.
Two deficiencies identified: incident resolution operating effectiveness; problem management escalation gap
5
Step 5 — Assess Severity and ReportThe 12 P1 SLA breaches on financially critical systems elevate this beyond a simple deficiency. If the affected systems include the general ledger or revenue recognition module, the likelihood and magnitude of potential misstatement may be sufficient to classify this as a significant deficiency in internal control over financial reporting. The auditor must communicate this to those charged with governance and may need to expand substantive testing to compensate for the ITGC weakness.
Potential significant deficiency — expanded substantive testing and governance communication required

Strengths, Limitations, and Framework Comparisons

Evaluating incident and problem management is not a one-size-fits-all exercise. CPAs encounter organizations that follow different frameworks, and each has distinct strengths and limitations. Understanding these trade-offs helps the evaluator calibrate expectations and tailor testing procedures to the organization's specific maturity level.

Comparison of major frameworks for incident and problem management
Framework / ApproachStrengthsLimitations
ITIL (v3/v4)Comprehensive, widely adopted, separates incident and problem management clearly, supports SLA-based measurement, extensive guidance on escalation and knowledge managementCan be overly bureaucratic for small organizations; implementation is expensive; does not directly map to financial audit assertions without interpretation
COBIT (DSS02/DSS03)Directly aligned with IT governance objectives; maps to COSO and SOX control requirements; provides maturity models for benchmarking; audit-oriented designLess prescriptive on implementation details; assumes the organization has chosen a complementary operational framework (e.g., ITIL)
ISO/IEC 20000International standard, certifiable, provides formal requirements (not just best practices), strong emphasis on documented proceduresCertification process is costly; smaller firms may find compliance burden disproportionate; less widespread in North American enterprises compared to ITIL
Ad Hoc / InformalLow overhead, flexible, may be sufficient for very small IT environments with limited complexityPoor auditability, inconsistent documentation, high person-dependency risk, difficult to demonstrate control effectiveness to external auditors
KEY TAKEAWAY
Frameworks like ITIL and COBIT are not competing standards—they are complementary layers, much like how GAAP provides accounting principles while internal audit provides the enforcement mechanism. COBIT tells the auditor what controls should exist; ITIL tells the IT department how to implement them. A CPA evaluator needs to understand both perspectives to perform an effective assessment.

Connection to Advanced Theory — IT Governance, Risk, and Assurance

Incident and problem management do not exist in isolation. They are embedded within a broader IT governance, risk, and compliance (GRC) ecosystem. Advanced topics that build upon the foundational concepts in this lesson include IT service continuity management (ITSCM), which addresses disaster recovery and business continuity at a strategic level; change management, which is the formal process through which permanent fixes identified by problem management are safely deployed into production; and configuration management, which maintains an accurate record of IT assets and their relationships, enabling faster incident diagnosis. For CPA candidates pursuing deeper expertise, understanding how these processes interact is essential for evaluating enterprise-level controls.

Mapping foundational concepts to advanced IT governance and assurance topics
This Lesson: Incident & Problem ManagementAdvanced Extension
Incident detection and loggingSecurity Information and Event Management (SIEM) for automated detection; integration with cybersecurity incident response plans
Root cause analysis in problem managementTrend analysis and predictive analytics using machine learning to anticipate incidents before they occur (AIOps)
Permanent fix via Request for Change (RFC)Formal change management processes (COBIT BAI06); integration with DevOps CI/CD pipelines for automated deployment
SLA-based performance measurementService level management and IT balanced scorecards aligned with enterprise KPIs and board-level risk reporting
Control deficiency classificationIntegrated assurance models combining ITGC testing with SOC 1/SOC 2 reporting for service organizations

As organizations increasingly adopt cloud-based financial systems, the locus of incident and problem management shifts from internal IT departments to third-party service providers. In such environments, the CPA must evaluate whether the organization obtains and reviews SOC 2 Type II reports from its cloud providers, which include detailed assessments of the provider's incident management controls. This evolution represents the frontier of the topic and is increasingly relevant on the CPA ISC examination.

Practice Problems

PROBLEM 1CONCEPTUAL
Explain the fundamental difference between incident management and problem management. Why is it important for a CPA evaluator to verify that an organization maintains both processes as distinct activities rather than combining them into a single workflow?
PROBLEM 2BASIC CALCULATION
An organization logged 800 incidents during the fiscal year. Of these, 680 were resolved within SLA targets, and 30 P1-Critical incidents out of 40 total P1 incidents met the one-hour SLA. Compute (a) the overall Incident Resolution Rate (IRR) and (b) the P1-Critical IRR. Identify which metric is more concerning from an audit perspective and explain why.
PROBLEM 3INTERMEDIATE
During your evaluation of a client's incident management process, you find that the organization has a well-documented incident management policy that includes a priority matrix, SLA targets, and escalation procedures. However, when you sample 50 incident records, you discover that 18 incidents were categorized as P4-Low despite involving the accounts receivable module during quarter-end processing. The incidents were resolved within the P4 SLA of 24 hours but would likely have warranted P2-High classification under the documented matrix. What type of control deficiency does this represent, and what are the audit implications?
PROBLEM 4APPLIED
Alpha Financial Services has migrated its core banking platform to a cloud service provider (CSP). Alpha's internal IT team no longer manages the infrastructure but remains responsible for application-level support. The CSP provides a SOC 2 Type II report that covers incident management controls. As the CPA evaluator, describe the steps you would take to evaluate incident and problem management in this hybrid environment. What specific elements would you look for in the SOC 2 report, and what complementary controls would you expect Alpha to maintain internally?
PROBLEM 5CRITICAL THINKING
A mid-size company reports that its Incident Resolution Rate has improved from 82% to 96% year-over-year, and its Mean Time to Resolve has decreased from 6.2 hours to 2.1 hours. Management presents these metrics as evidence that incident management controls are operating effectively. As a skeptical CPA evaluator, identify at least three reasons why these improved metrics might not actually indicate effective controls. What additional evidence would you request to corroborate the metrics?

Lesson Summary

Evaluating incident management and problem management is a core competency for CPA candidates in the ISC domain. Incident management focuses on rapid service restoration through detection, logging, categorization via the impact × urgency priority matrix, escalation, and resolution within SLA targets. Problem management complements this by performing root cause analysis to eliminate recurring incidents and maintain the Known Error Database (KEDB). Key quantitative metrics—the Incident Resolution Rate, Mean Time to Resolve, and Recurrence Rate—provide auditable evidence of control effectiveness.

The CPA evaluator distinguishes between deficiencies in design (missing policies or matrices) and deficiencies in operating effectiveness (policies that exist but are not consistently followed). Frameworks such as ITIL and COBIT provide complementary lenses—COBIT defines control objectives aligned with SOX and COSO, while ITIL provides implementation guidance. In modern cloud environments, the evaluation extends to reviewing SOC 2 Type II reports from service providers and verifying that complementary user entity controls are in place. The ultimate objective is to ensure that IT disruptions do not compromise the integrity of financial reporting.

Varsity Tutors • CPA (ISC) • Evaluate Incident And Problem Management