An AI incident response plan prepares an organization to recognize unexpected behavior, stop ongoing harm, preserve evidence, coordinate decisions, correct downstream effects and restore service safely. It should connect AI-specific events to existing security, privacy, legal, operational and crisis processes.
Use the tracker above during preparation and response, but do not let a checklist delay emergency action or required reporting. Severity is provisional: escalate whenever impact, affected data, autonomy, spread or uncertainty increases.
1. Purpose of an AI incident response plan
The plan reduces confusion when an AI-enabled workflow fails or behaves outside approved expectations. It establishes authority, communication, evidence, containment options and recovery checks before pressure is high.
NIST finalized SP 800-61 Revision 3 in April 2025 to integrate incident response across cybersecurity risk management. Its emphasis on preparation, detection, response and recovery is useful here, but AI incidents can also involve inaccurate or unfair output, harmful decisions, intellectual-property issues or unexpected autonomy without a conventional cyber intrusion.
2. Define which events activate the plan
Include suspected data exposure, unauthorized access, prompt injection, harmful or discriminatory output, material inaccuracy, unsafe automated action, provider compromise, model or configuration change, loss of required records, unapproved tool use and failure of human oversight.
Define thresholds for an event, incident and crisis. Record affected people, data, decisions, systems, customers, jurisdictions and dependencies. Keep a route for employees to report uncertainty without having to prove root cause.
3. Assign an incident lead and specialist roles
The incident lead coordinates severity, priorities, owners and closure. Technical owners investigate configuration, models, accounts, logs and integrations. Operational owners identify affected work and safe alternatives. Security, privacy, legal, compliance, communications, HR, procurement, records and vendor contacts join according to the event.
Pre-authorize emergency actions where appropriate: pausing automation, disabling a token, isolating an integration, switching to manual work or preventing publication. Document who may approve restoration.
Prepare before an incident occurs
Maintain current contacts, system and data-flow diagrams, approved-tool inventory, logging locations, access procedures, safe manual alternatives and vendor escalation routes. Confirm that responders can obtain necessary evidence without relying on the unavailable or compromised service.
Run tabletop exercises using realistic scenarios such as a meeting bot sharing restricted content, an agent taking an unauthorized action or a model update changing output quality. CISA provides tabletop exercise packages that illustrate how structured scenarios can test roles, decisions and communications before a real event.
4. Triage severity from current evidence
| Level | Working description | Response posture |
|---|---|---|
| 1 — Low | Contained, minor and readily reversible | Owner handles and records |
| 2 — Moderate | Limited data, people or workflow affected | Coordinate relevant specialists |
| 3 — High | Material impact, sensitive data or uncertain spread | Incident lead and executive escalation |
| 4 — Critical | Ongoing severe harm, broad compromise or major dependency | Crisis procedures and urgent external support as needed |
Consider impact, scope, sensitivity, reversibility, autonomy, affected people, legal or contractual triggers, public exposure and confidence in containment. Reclassify as evidence changes.
5. Contain without destroying evidence
Pause the risky action, not necessarily every service. Disable or restrict affected automation, credentials, tokens, accounts or integrations according to technical and operational advice. Protect connected systems and provide a safe manual alternative.
Preserve relevant prompts, outputs, logs, timestamps, model and version, settings, user actions, access, data flows and decisions. Do not collect or copy more sensitive information than the investigation requires. Maintain integrity, access control and a timeline.
Do not “fix first and document later” when the change would erase the only evidence. Balance immediate harm reduction with preservation, involving qualified responders for high-impact events.
6. Investigate cause and downstream impact
Distinguish the trigger, root cause and contributing conditions. Possible causes include permissions, unsafe defaults, model behavior, poor test coverage, misleading input, adversarial activity, integration logic, provider change, missing review or pressure that caused users to bypass controls.
Find every downstream use of affected output: messages, decisions, files, customer records, code, reports and automated actions. Correcting the source does not repair what has already propagated.
Coordinate with the provider without surrendering ownership
Use the vendor’s security or support channel, record case identifiers and preserve responses. Ask for relevant timestamps, scope, model or service changes, containment status and recovery evidence. Do not wait for the provider to classify organizational impact or decide the organization’s notification obligations.
Keep an independent timeline and decision log. A provider may investigate infrastructure while the customer must address affected people, records, business continuity and downstream actions.
7. Coordinate notification, recovery and learning
Assess contractual, legal, regulatory, insurer, customer, employee, vendor and law-enforcement obligations with appropriate specialists. Communicate confirmed facts, uncertainty, protective actions and the next update. Avoid speculative root-cause claims.
Recover in stages. Verify accounts, permissions, integrations, data handling, output quality, human review and monitoring before full restoration. Require explicit approval from the incident lead or authorized owner.
Conduct a blameless review covering impact, timeline, decisions, evidence gaps and control changes. Assign actions, owners and dates. Update the AI acceptable use policy, risk assessment, test cases and training.
8. Incident examples
Sensitive meeting shared incorrectly
Pause sharing, restrict access, preserve audit information, identify recipients and content, coordinate privacy and security review, correct copies where possible and verify workspace settings before restoration.
Automated support action sends wrong information
Stop the automation, switch to a reviewed process, find affected customers and records, preserve the inputs and rules, correct downstream information and test the redesigned workflow.
Prompt injection through connected content
Isolate the integration, protect credentials and connected systems, preserve malicious input and tool activity, inspect the full action chain and reduce permissions before controlled testing.
Common incident-response mistakes
- Waiting for proof of root cause before containing ongoing harm.
- Deleting prompts, logs or accounts before preserving relevant evidence.
- Focusing on the model and missing permissions or downstream automation.
- Leaving operational, privacy or affected-person impact outside the response.
- Using severity as fixed despite new evidence.
- Restoring the same workflow without testing the corrective control.
- Closing the incident without assigned follow-up actions.
AI incident response plan FAQ
Is every wrong AI answer an incident?
No. Use defined thresholds based on impact, spread, data, decisions and control failure. Repeated minor errors may still reveal a systemic issue requiring escalation.
Should the AI tool be shut down immediately?
Contain the risky path proportionately. Broad shutdown may be necessary for severe or uncertain harm, but targeted restriction can preserve safe operations and evidence.
What evidence should be preserved?
Relevant prompts, outputs, logs, versions, settings, access, integrations, timestamps, decisions and downstream use—collected lawfully and with appropriate access and integrity.
When can normal use resume?
After the cause and impact are sufficiently understood, corrective controls are tested, critical obligations are addressed and the authorized incident lead approves staged recovery.
Methodology and limitations
This AI incident response plan adapts established prepare, detect, respond and recover concepts to AI-enabled workflows, adding output, autonomy, affected-person and downstream-decision considerations. The tracker stores progress only in the browser.
The guide is general information and not emergency, legal, cybersecurity, privacy or regulatory advice. Use existing organizational plans and qualified specialists, and comply with applicable notification and evidence requirements.