1. What you will learn
Good governance starts with an honest picture of how AI systems fail. This lesson gives you a structured failure-mode map so you can diagnose risks in any AI use in your business. You will learn to:
- classify AI failures into data, model, deployment and human failure modes;
- explain common mechanisms such as unrepresentative data, proxy discrimination, overfitting, automation bias and feedback loops;
- link each failure mode to the harm it can cause;
- identify the control that addresses each failure mode;
- run a quick diagnostic on one of your own AI uses.
2. The idea explained
An AI system fails when its outputs, or the actions taken on them, cause unacceptable harm or fail to deliver the intended benefit. Failures rarely come from one cause. It helps to trace them to four layers.
A. Data failure modes
- Unrepresentative data — the training data does not reflect the people or situations the system will meet. A speech tool trained mainly on one accent performs worse on others; a model trained on metro customers misjudges rural ones.
- Historical bias — data faithfully records past decisions that were themselves unfair, and the model learns to repeat them.
- Proxy variables — a seemingly neutral variable stands in for a protected or sensitive characteristic, such as PIN code, first name or college attended acting as proxies for community, gender or class.
- Label errors — the "correct answers" in training data are wrong or inconsistent, so the model learns noise.
- Data leakage (in modelling) — information that would not be available at decision time sneaks into training, making test results look better than live performance.
- Unlawful or unconsented data — the business had no right to use the data for this purpose.
B. Model failure modes
- Overfitting — the model memorises the training data and generalises poorly to new cases.
- Hallucination — generative models produce fluent but false content, including invented facts, figures, quotations or references.
- Opacity — the model cannot give meaningful reasons, making it hard to detect errors or explain decisions.
- Brittleness — small, unusual changes in input (spelling, formatting, a new product name) produce large errors.
- Vulnerability to attack — prompt injection, data poisoning or adversarial inputs manipulate outputs.
C. Deployment failure modes
- Purpose creep — a tool built for one purpose is used for another without re-assessment.
- Drift — the world changes and performance quietly declines.
- Wrong threshold — a cut-off score set without considering the cost of each type of error.
- Feedback loops — the system's outputs shape the future data it learns from. A policing or fraud model that sends more checks to one area finds more cases there, "confirming" its own prediction.
- Integration errors — mistakes in how outputs are passed to other systems, such as a probability read as a yes or no.
D. Human and organisational failure modes
- Automation bias — people over-trust machine outputs and stop applying judgement, especially under time pressure.
- Algorithm aversion — the opposite: people ignore a reliable tool, losing its benefit.
- Unclear accountability — no one owns the system, so no one notices or acts on problems.
- Inadequate training — users do not know the tool's limits.
- Incentive misalignment — targets reward speed or volume, encouraging staff to accept outputs without checks.
From failure mode to harm. Each mode maps to harms introduced in Start here 01: accuracy, fairness, privacy, security, legal and reputational harm. Unrepresentative data often causes fairness harm; hallucination causes accuracy and legal harm; purpose creep often causes privacy harm; automation bias magnifies every other failure.
From failure mode to control. Controls are matched to modes:
- data audits and representativeness checks for data failures;
- validation on held-out, realistic data and group-wise error analysis for model failures;
- grounding generative models in approved sources and requiring verification for hallucination;
- purpose limitation, change control and monitoring for deployment failures;
- meaningful human review, training and balanced incentives for human failures.
3. Let us work through it
Step 1 — Describe the use. One sentence on purpose, inputs, output and who acts on it.
Step 2 — Walk the four layers. For each of data, model, deployment and people, ask "what could go wrong here?" and list the relevant modes.
Step 3 — Link to harm. For each mode, write the harm and who suffers it.
Step 4 — Rate roughly. Mark each as likely or unlikely, and serious or minor. The full scoring method comes in Reader IV.
Step 5 — Name a control. Match at least one control to each likely or serious mode.
Step 6 — Record in the inventory. Add the top three failure modes and controls to the use's entry.
Worked example
4. Worked examples
Example 1 — CV screening. A company trains a screening model on ten years of hiring decisions in which few women were hired for field-sales roles. Data layer: historical bias; proxies such as participation in certain sports or colleges. Model: opaque scores. Human: recruiters under volume pressure accept scores (automation bias). Harm: fairness and legal risk, lost talent. Controls: remove proxy features after analysis, test selection rates by gender, require recruiters to review a sample of rejected CVs, and set a policy that the tool ranks but never rejects alone.
Example 2 — Customer-service generative assistant. A bank-partner fintech's assistant tells a customer a charge will be refunded when policy does not allow it. Model layer: hallucination. Deployment: no grounding in the fee schedule. Harm: accuracy, consumer protection and reputational harm. Controls: retrieval from the approved policy documents only, refusal when the answer is not in the documents, and escalation to staff for money-related commitments.
Example 3 — Inventory forecasting with a threshold problem. A pharmacy chain's reorder model outputs a probability of stock-out. Integration code treats anything above 0.5 as "reorder", but for critical medicines the cost of a stock-out is far higher than the cost of excess stock. Deployment failure: wrong threshold. Control: set item-specific thresholds based on cost of each error, with lower thresholds for critical items.
Example 4 — Feedback loop in collections. A collections model prioritises calls to customers it predicts will default. Staff call them more often; some complain of harassment; complaint data then marks them as "difficult", raising their risk scores further. Controls: cap contact frequency in line with applicable fair-practice rules, separate complaint data from risk features, and review outcomes by customer segment quarterly.
5. Common mistakes and how to fix them
- Blaming "the algorithm" for every failure. Fix: trace the failure through data, model, deployment and people.
- Testing only overall accuracy. Fix: analyse errors by group and by situation.
- Setting thresholds by default. Fix: set cut-offs using the cost of each type of error.
- Ignoring automation bias. Fix: design review so people must record reasons when they accept or override.
- Letting tools drift into new purposes. Fix: require re-assessment for every new purpose.
- Assuming generative tools know your policies. Fix: ground them in approved documents and require verification.
Key takeaways
6. Board summary
AI fails in four layers: data, model, deployment, people. Bias comes from unrepresentative data, historical decisions and proxies. Generative AI hallucinates; ground it and verify it. Thresholds must reflect the cost of each error. Automation bias magnifies every other failure. Match each likely or serious failure mode to a named control.
Check your understanding
7. Practice and self-check
- Name the four layers of AI failure.
Answer: Data, model, deployment, and human or organisational.
- What is a proxy variable?
Answer: A seemingly neutral variable that stands in for a sensitive characteristic, such as PIN code for community.
- Define automation bias.
Answer: The tendency of people to over-trust machine outputs and stop applying their own judgement.
- What is a feedback loop?
Answer: When a system's outputs shape the future data it learns from, reinforcing its own predictions.
- Which failure mode is purpose creep?
Answer: A deployment failure, where a tool is used for a new purpose without re-assessment.
- What control addresses hallucination in a customer assistant?
Answer: Grounding answers in approved documents, refusing when unsupported, and escalating commitments to staff.
- Why might test accuracy exceed live accuracy?
Answer: Because of overfitting or data leakage in modelling.
- How should a threshold be set?
Answer: By comparing the cost of false positives and false negatives for that use.
- Which harm does historical bias usually cause?
Answer: Fairness harm, often with legal and reputational consequences.
- What is algorithm aversion?
Answer: Ignoring a reliable tool's outputs, losing its benefit.