The assessment examines the capabilities needed to understand AI dependencies, maintain essential work during disruption, manage cost pressure, reduce supplier concentration and improve incident recovery. Each group below contains three capabilities, creating a consistent fifteen-capability view of AI-related business continuity.
AI Dependency Mapping and Risk Assessment
This group is about knowing exactly where your organisation leans on AI and how much trouble you are in if it stops. It covers listing the AI systems you truly depend on, finding hidden points where one failure spreads widely, and judging the real operational and financial impact of an outage.
- Critical AI System Identification
- Single Point of Failure Analysis
- Business Impact Assessment for AI Outages
This is about having a clear, up-to-date picture of which AI systems your organisation actually relies on to function. It means listing every model, tool, and service in use, working out which ones are essential versus nice-to-have, and ranking them by how much damage their loss would cause. Without this, you cannot protect what matters most. This looks for the places where one AI tool, model, provider, or shared component could fail and take down a whole process or many processes at once. It means tracing how your systems connect, spotting where everything funnels through a single dependency, and flagging those choke points before they cause a much wider, damaging collapse. This is about understanding, in practical terms, what happens when a key AI system goes down. It means knowing which business processes stop, how fast the disruption spreads to customers and staff, how long you could cope, and roughly what each hour of downtime costs. This turns vague worry into clear, prioritised planning.
Resilience and Failover Planning
This group covers what keeps you running when AI is unavailable or unreliable. It looks at whether staff can fall back to manual or simpler methods, whether you have backup models and providers ready to take over, and whether your services can scale back safely instead of failing all at once.
- Manual Fallback and Workaround Procedures
- Redundancy and Alternative Provider Readiness
- Graceful Service Degradation Design
This is about having documented ways for people to keep essential work going when the AI is simply not there. It means writing down the manual steps, simpler tools, or older methods staff should switch to, making sure people know about them, and keeping them usable so nobody is left frozen during an outage. This covers having more than one way to get the AI capability you need. It means setting up backup models, second providers, or alternative regions in advance, testing that you can actually switch to them, and being able to reroute work quickly when your primary system fails, slows down, or becomes unavailable. This is about building services that scale back safely when AI capacity drops, rather than crashing completely. It means deciding in advance which features to switch off first, allowing slower or simpler responses under strain, and making sure that losing the AI reduces quality temporarily instead of stopping the whole service for everyone.
AI Cost and Usage Continuity
This group is about making sure rising or unpredictable AI costs never force you to stop essential work. It covers watching and forecasting your token and usage spend, throttling low-value activity to protect critical tasks, and holding financial buffers so a sudden price or demand shock does not break operations.
- Token Cost Monitoring and Forecasting
- Usage Throttling and Prioritisation Controls
- Budget Contingency and Shock Absorption
This is about seeing your AI spending clearly and early, rather than discovering it on the invoice. It means tracking token and usage costs close to real time, breaking them down by team, product, or feature, and forecasting where spend is heading so you can act on worrying trends well before they become emergencies. This covers your ability to control AI usage deliberately when costs climb or capacity tightens. It means being able to slow down or pause low-value activity, set limits per team or task, and steer scarce capacity toward the work that matters most, all without having to shut the entire AI service down at once. This is about having financial room to handle sudden AI cost rises without halting essential work. It means building buffers into budgets, agreeing what gets funded first when prices jump or usage surges, and planning for provider price changes in advance, so a cost shock becomes a managed decision rather than a forced shutdown.
Vendor and Supply Chain Resilience
This group focuses on the outside providers your AI depends on and the risk they carry. It covers understanding each vendor's reliability and service promises, keeping your data and models portable so you are not trapped, and having a tested plan to switch providers if one fails or becomes too expensive.
- AI Vendor Risk and SLA Management
- Data and Model Portability
- Vendor Exit and Substitution Planning
This is about knowing how dependable your AI vendors really are and holding them to clear commitments. It means assessing each provider's reliability, security, and financial health, understanding the service levels they promise, and making sure your agreements give you real protection, notice, and recourse when they have outages or change their terms. This covers how easily you could move away from a given AI provider or model if you needed to. It means keeping your data, prompts, and configurations in open, exportable formats, avoiding designs that lock you into one supplier, and making sure that switching would be a manageable project rather than an impossible one. This is about being ready to replace an AI provider that fails, raises prices sharply, or shuts down. It means knowing which alternatives exist, keeping a tested plan for moving to them, and understanding the time, cost, and effort a switch would take, so a vendor problem never leaves you completely stranded and helpless.
Incident Response, Recovery and Learning
This group is about being ready for AI disruptions and getting better after each one. It covers rehearsing outage and cost scenarios so plans are proven, detecting and responding to failures quickly with clear roles, and reviewing every incident honestly so the same problem does not keep coming back to hurt you.
- AI Continuity Testing and Drills
- Incident Detection and Response
- Post-Incident Review and Improvement
This is about proving your continuity plans work before a real crisis tests them. It means regularly rehearsing scenarios like an AI outage or a sudden cost spike, walking through your fallbacks and responses, and finding the gaps in a safe practice setting so your plans are trusted and proven rather than just written down. This covers spotting AI failures quickly and reacting in an organised way. It means monitoring for outages, errors, and degraded performance, having clear alerts, knowing exactly who does what when something goes wrong, and following agreed steps to contain the impact, communicate with people affected, and restore normal service as fast as is safely possible. This is about learning honestly from every AI disruption so the same thing does not keep happening. It means reviewing what went wrong without blame, capturing the lessons, and actually updating your plans, controls, and training as a result, turning each incident into a real improvement rather than a problem you quietly repeat later.