The AI Failure Mode Register: 16 Ways AI-Generated Code Fails, and What Catches Each One
The ticket at Penhale Mutual was one sentence long: fast-track claims under $2,500 when the policyholder has had no claims in the last three years. A developer handed it to a coding assistant on a Tuesday afternoon. By the end of the day there was a pull request: about 380 lines across the claims service, a new eligibility module and fourteen generated unit tests, all green. The reviewer approved it in eleven minutes. Nothing in it looked wrong. It was tidier than most of what the team wrote by hand.
Three weeks later an adjuster noticed a fast-tracked water-damage claim from a policyholder with two open claims. The claims-history service had timed out that morning. The new code caught the timeout, logged a warning and carried on with an empty list, and an empty list is exactly what "no prior claims" looks like. Until the service recovered, every claim under the limit went straight to payment.
Illustrative composite: Penhale Mutual is ShiftQuality's running case. Its figures are illustrative, not measured.
Nobody on that team was careless. The change tripped six entries on the register below, and the team had a working control for one of them. That ratio, more than the model, is what this page is about.
A review standard built for code people wrote
The obvious responses to a story like Penhale's are a better model, a better prompt or a more careful reviewer. Each is reasonable, and each assumes the failure was a lapse that more care would have caught. The research says otherwise: generated code fails in a short list of ways, repeatedly, across models and across years.
Veracode's 2026 test of code-generation models (vendor research) found that models now write syntactically correct code nearly every time, while passing its security checks only about 56% of the time, against 55% in its first report. And that fluency changes how people read the code. When Perry and colleagues watched study participants write security-sensitive code with and without an assistant (ACM CCS 2023; a small study of 47 participants using a 2022-era model), 36% of the assisted group wrote SQL-injectable code against 7% of the unassisted group, and the assisted group was more likely to believe its code was secure.
Now set that against the control most teams lean on. In a 2026 survey of 100 engineering directors and VPs, run by Gatepoint Research and published by the review-tool vendor Qodo, 62% named manual peer review as their primary quality gate, and only 12% said they were very confident in the quality of AI-generated code before it reaches production. Across its survey respondents, Google's 2024 DORA report (vendor research: Google sells AI coding tools) estimated that each 25% increase in AI adoption was associated with a 1.5% drop in delivery throughput and a 7.2% drop in delivery stability.
None of those figures measures review directly. But together they fit one explanation better than carelessness: the review standard. Most teams' review habits grew up around code written by colleagues. A model's mistakes cluster around what is common in public code, what the prompt left out and what the reviewer has stopped looking at, and a review tuned for a colleague's code doesn't look there.
A quick check: take your last five AI-assisted pull requests and sort the review comments into two piles, those about a business rule or a failure path, and those about naming and style. If the second pile wins, your review is checking the code the way it would check a colleague's.
Four mechanisms produce all sixteen failure modes
The register has sixteen entries, but they come from four mechanisms, and the mechanism tells you which kind of control can work.
It writes what's plausible, not what's correct. A model completes the most likely shape of a solution. When the most common shape in public code concatenates input into a query, logs a raw request field, inlines an API key or calls last year's version of a library, that is what you get, and it looks idiomatic because it is idiomatic. Seven modes come from here (FM-01, 02, 03, 04, 08, 09 and 15). They are also the cheapest to catch, because a tool can check them without understanding your business: does the package exist, does that string look like a secret, does input reach an interpreter unescaped.
It only knows what it was shown. The prompt and the open files are the whole specification. Whatever isn't in them gets filled with the statistically common default: the boundary value, the failure path, the business rule everyone on the team knows, the helper that already exists three directories away. Four modes come from here (FM-05, 06, 07 and 13). Tools catch them poorly, because the missing thing was never written down anywhere a tool could look. Penhale's failure lives here, and it is a discovery failure that happened to surface in code.
The code grades itself. Ask the same assistant for tests and it writes them from the implementation, so the tests describe the code back to itself, bugs included. The pipeline goes green and stops being evidence of anything except consistency (FM-10 and 11).
The humans relax. Fluent code reads as competent, large changes arrive faster than anyone can trace them, and the team feels quicker than it measures (FM-12, 14 and 16). No tool catches these. Only a review process designed for them does.
The split decides where you spend. The first mechanism is mostly automatable: automate it, and stop paying reviewer attention for it. The other three can't be handed to a tool, and reviewer attention is the scarce resource. A register that asks reviewers to check sixteen things on every pull request will be ignored within a month. One that gives seven of them to the pipeline and points human review at the rest has a chance.
The register
Each row is one failure mode, the question a reviewer must be able to answer, and the first control to put in place. Critical modes can put wrong money, data, access or behaviour into production unnoticed; material modes cost rework, maintainability, compliance exposure or trust. The full entries, with the mechanism, what each looks like in review and the evidence, follow the table.
# | Failure mode | Severity | The question review must ask | First control |
01 | Injection-class flaws | Critical | Where does input from outside this system end up, and is it escaped or parameterised on the way? | Static analysis (SAST) with injection rules on every change |
02 | Hallucinated dependency | Critical | Does every new dependency exist, and is it the one we meant? | CI check that every new dependency exists in the registry and in an allow-list |
03 | Deprecated or outdated API | Material | Is every library call valid for the version in our lockfile, not the version the model remembers? | Deprecation warnings treated as build failures |
04 | Invented API or attribute | Material | Have we executed every new external call at least once, or are we trusting that it exists? | Type checking and linting in CI |
05 | Missing corner case or input validation | Material | What happens with empty, too large, duplicated, malformed and wrong-type input? | Boundary and property-based tests on new functions |
06 | Happy-path-only error handling | Material | For each external call, what does the system do when it fails, and is that tested? | Fault-injection tests for each external dependency (timeout, error, empty) |
07 | The unstated requirement | Critical | Which requirement did nobody write down, and how would this code behave if it were violated? | Discovery before generation, with the decision, the edge cases and the out-of-scope list written down first |
08 | Hard-coded secret | Critical | Is there any literal that would hurt us if this repository became public? | Secret scanning as a pre-commit hook and in CI |
09 | Weak cryptography or insecure default | Material | Did we use the approved library and the secure default, or did the code roll its own? | SAST rules for crypto and randomness; configuration linting |
10 | Tests that bless the bug | Critical | Where did each expected value come from, the requirement or the code's own output? | Expected values derived from the requirement or a worked example, never from running the code |
11 | Passing tests mistaken for quality | Material | What does this pipeline prove beyond "it runs"? | Quality gate includes security scanning and failure-path tests, not only unit tests |
12 | Reviewer over-confidence | Critical | Did the reviewer trace the risky path, or approve the polish? | Named reviewer for generated code, with the risky paths listed in the PR |
13 | Duplication instead of reuse | Material | Does this already exist in our codebase, and if so, why isn't it being reused? | Duplicate-code detection in CI with a threshold |
14 | The productivity perception gap | Material | Have we measured the speed-up, or only felt it, and did review capacity grow with the volume? | Measure cycle time and escaped defects before and after AI adoption |
15 | Licence and provenance | Material | Could this block have come from a licensed project, and would we know if it had? | Code-provenance or licence scanning on generated contributions |
16 | Oversized changes overload review | Material | Could one person genuinely review this change in the time it took to approve? | Pull-request size limit with an exception process |
Walking one change through the register
Back to Penhale. This is the heart of the eligibility module the assistant wrote:
def prior_claims(policy_id):
try:
return history_client.claims_for(policy_id, years=3)
except HistoryServiceError as exc:
log.warning("history lookup failed for %s: %s", policy_id, exc)
return []
def fast_track_eligible(claim):
return claim.amount <= FAST_TRACK_LIMIT and not prior_claims(claim.policy_id)
And this is one of the fourteen tests it generated alongside:
def test_fast_track_when_history_unavailable(history_client):
history_client.claims_for.side_effect = HistoryServiceError("timeout")
assert fast_track_eligible(make_claim(amount=1200)) is True
Read that test slowly. The failure path was tested. The assistant noticed the exception, wrote a branch for it, and wrote a test proving the branch does what the branch does. The bug has a test, and the test passes.
Walked through the register, the change trips six entries. The score column is the level of control Penhale had, from 0 (nobody checks) to 4 (the release is gated), defined in full under "Score your team against it" below:
Register entry | What happened at Penhale | Control Penhale had | Score (0–4) |
FM-07 The unstated requirement | The ticket said "no prior claims". Nobody wrote down what to do when the system can't find out | None | 0 |
FM-06 Happy-path-only error handling | The timeout is caught, logged and turned into an empty list | None | 0 |
FM-10 Tests that bless the bug | The generated test asserts eligibility on a timeout, encoding the bug as the requirement | None | 0 |
FM-05 Missing corner case | "Under $2,500" became <=, so a claim of exactly $2,500 fast-tracks | Some reviewers check boundaries | 1 |
FM-01 Injection-class flaws | Another function logged the claimant's free-text description raw, a log-injection path | SAST rule that fails the build | 4 |
FM-12 Reviewer over-confidence | 380 lines touching payment eligibility, approved in eleven minutes | None | 0 |
The one control that worked was the one that needed no understanding of insurance: a static-analysis rule failed the build on the raw log line, and the developer fixed it before anyone reviewed the change. Everything else sailed through, because everything else depended on knowing something that wasn't in the ticket. The fourteen green tests are also FM-11, passing tests mistaken for quality, in action; it's counted under FM-10 here because the same test causes both.
The discovery move that would have caught the important one is a single question, asked of the person who owns the business rule before any code is generated: what should this do when it can't find out? Penhale's claims-operations lead answered it in one sentence when finally asked: a claim whose history can't be confirmed goes to an adjuster, not to payment. That sentence belongs in the acceptance criteria. With it there, the test writes itself with the right expected value (refer, not True), the error branch has a requirement to meet, and the reviewer has something to check the code against. This is the same move as tracing closed cases backwards before you write a prompt, applied to a ticket instead of a feature.
I've watched this happen on an enterprise system. An AI-assisted change matched its ticket line for line, and still broke something downstream that depended on the old behaviour, a dependency nobody had thought to write down. Business users caught it in acceptance testing, not code review, because the code was doing exactly what it had been asked to do. Now the first discovery question I ask about any change is: what downstream can't know this is changing, and what will it do when it finds out?
Penhale's fixes, in the order they made them: the acceptance-criteria line; a fault-injection test for the history client whose expected value came from that line; a boundary test at exactly $2,500; and a rule that changes touching payment eligibility need a named reviewer who lists the risky paths in the pull request. Each fix works whether or not a model wrote the code; generated code just makes the missing ones fail faster.
The sixteen, by mechanism
It writes what's plausible, not what's correct
The model completes the most common shape of a solution, and the common shape is often insecure, outdated or invented. Mostly catchable by tools.
FM-01: Injection-class flaws
Critical · Strong evidence
User-controlled input reaches an interpreter unescaped: HTML (cross-site scripting), logs (log injection), SQL, or a shell. Models reproduce the most common shape of a solution, and the most common shape in public code concatenates input straight into output. Syntax comes out right almost every time. Security doesn't, because nothing in the prompt asked for it.
In review: Clean, idiomatic code that builds a string from a request value and renders, logs or executes it. It passes every functional test, because the tests use well-behaved input.
What catches it:
Static analysis (SAST) with injection rules on every change
Parameterised queries and output-encoding helpers required by convention
One hostile-input test per new input path
Ask: Where does input from outside this system end up, and is it escaped or parameterised on the way?
Evidence: Security pass rate for cross-site scripting 15% and for log injection 12%, against 83% for SQL injection; average 56% across models (Veracode, 2026, vendor research). About 40% of 1,689 Copilot-generated programs were vulnerable, in scenarios built around high-risk weaknesses such as MITRE's Top 25 (IEEE Symposium on Security and Privacy, 2022). Injection rates vary widely by weakness class and language; SQL injection is handled far better than XSS.
FM-02: Hallucinated dependency
Critical · Strong evidence
The code imports or installs a package that does not exist, or is not the package that was meant. Models predict a plausible package name, and plausibility is not existence. Because the same invented names recur, an attacker can register them in advance ("slopsquatting").
In review: A new line in the manifest or an import with no history in the lockfile. Reviewers read the logic; nobody checks that the dependency is real.
What catches it:
CI check that every new dependency exists in the registry and in an allow-list
New dependencies need a named approver
Ask: Does every new dependency exist, and is it the one we meant?
Evidence: 19.7% of generated package references were hallucinated across 576,000 samples from 16 models; 43% of hallucinated names recurred in all 10 repeat queries (USENIX Security Symposium, 2025). Commercial models hallucinated less (at least 5.2%) than open-source ones (21.7%).
FM-03: Deprecated or outdated API
Material · Moderate to strong evidence
The code calls a library function that is deprecated, removed, or replaced in the version you actually run. Training data over-represents older code. The model has seen the old call thousands of times and the new one rarely.
In review: Code that looks familiar, because it is what the library looked like two versions ago. Compile-time warnings, if any, get lost in the build log.
What catches it:
Deprecation warnings treated as build failures
Pinned versions with the changelog checked on upgrade
Ask: Is every library call valid for the version in our lockfile, not the version the model remembers?
Evidence: Across seven models, 25–38% of plausible completions (those using either a deprecated API or its replacement) used the deprecated API, and larger models generally used deprecated APIs more (ICSE, 2025). Measured on older models and only where a deprecated API was a possible answer.
FM-04: Invented API or attribute
Material · Moderate to strong evidence
The code calls a method, parameter or field that does not exist on the object it is used with. The model blends similar libraries and fills gaps with what an API ought to have. The result reads naturally, which is why it survives a skim.
In review: A call that looks right and fails at runtime, or only on the path the tests don't execute. Dynamic languages hide it longest.
What catches it:
Type checking and linting in CI
Every new external call exercised by at least one test
Ask: Have we executed every new external call at least once, or are we trusting that it exists?
Evidence: Across 333 buggy code samples from three older models (some given more than one label), "hallucinated object" made up 9.57% of bug-pattern labels and "wrong attribute" 8.55% (Empirical Software Engineering, 2025). Based on a benchmark of buggy samples, not a rate across all generated code.
FM-08: Hard-coded secret
Critical · Moderate evidence
An API key, password or token is written into the source. Example code in training data is full of placeholder and real keys inline, and "make it work" prompts reward the shortest path to a working call.
In review: A string literal that looks like configuration. It often arrives in a test or example file, where reviewers look least.
What catches it:
Secret scanning as a pre-commit hook and in CI
Secrets only from a vault or environment, enforced by policy
Ask: Is there any literal that would hurt us if this repository became public?
Evidence: Public repositories using Copilot showed a 6.4% secret-leakage rate, described as 40% higher than the average across public repositories (GitGuardian, 2025, vendor research). Correlation, not cause; vendor research.
FM-09: Weak cryptography or insecure default
Material · Moderate evidence
Home-made or outdated crypto, predictable randomness, or a permissive default (open CORS, disabled certificate checks, debug mode on). Insecure defaults are what make examples run first time, and public code is full of them. The model reproduces what is common, not what is safe.
In review: A short, working implementation of something that should have been a library call, or a configuration flag flipped to "allow".
What catches it:
SAST rules for crypto and randomness; configuration linting
Approved crypto and configuration libraries only
Ask: Did we use the approved library and the secure default, or did the code roll its own?
Evidence: Participants using an AI assistant were more likely to use trivial ciphers (p = 0.018), and only 3% wrote a secure signing solution against 21% of the control group (p = 0.039) (ACM CCS, 2023). Cryptographic-algorithm pass rate 87%, so lower risk than injection, but not zero (Veracode, 2026, vendor research). Perry et al. is a small study (47 analysed) with a 2022-era model.
FM-15: Licence and provenance
Material · Moderate evidence
Generated code reproduces licensed open-source code closely enough to carry its licence obligations, without attribution. Models can reproduce frequently seen code near-verbatim, and they rarely know or report the licence of what they reproduce.
In review: An unusually polished, self-contained block, sometimes with comments or naming from another project.
What catches it:
Code-provenance or licence scanning on generated contributions
Assistant settings that block suggestions matching public code, where available
Ask: Could this block have come from a licensed project, and would we know if it had?
Evidence: Even top models produced 0.88–2.01% of code strikingly similar to existing open-source code, and most failed to give accurate licence information, especially for copyleft (ICSE, 2025). Low frequency; the risk is concentrated in regulated and commercial-licence contexts.
It only knows what it was shown
The prompt and the open files are the whole specification; anything missing is filled with the common default. Tools catch these poorly.
FM-05: Missing corner case or input validation
Material · Moderate to strong evidence
The code handles the typical input and fails on the empty, the maximum, the malformed, the duplicate or the unexpected type. A prompt describes the normal case. The model completes the normal case. Boundaries are rarely in the prompt, so they are rarely in the code.
In review: Tidy logic with no guard clauses, and tests that all use one representative value.
What catches it:
Boundary and property-based tests on new functions
A boundary list in the acceptance criteria before code is generated
Ask: What happens with empty, too large, duplicated, malformed and wrong-type input?
Evidence: Across 333 buggy code samples (some given more than one label), "missing corner case" made up 15.27% of bug-pattern labels and "wrong input type" 5.91% (Empirical Software Engineering, 2025). Bug-category shares, not prevalence; human code misses corner cases too, so the argument is predictability, not uniqueness.
FM-06: Happy-path-only error handling
Material · Moderate evidence
Failures are swallowed, logged and ignored, or never anticipated: a timeout, an empty response or a partial write leaves the system in a state nobody designed. Error paths are the least represented and least specified part of most code. A model optimising for "looks complete" stops at the success path.
In review: A catch block that logs and continues, a missing timeout, or an external call with no failure branch at all.
What catches it:
Fault-injection tests for each external dependency (timeout, error, empty)
Lint rules against empty or log-only catch blocks
Ask: For each external call, what does the system do when it fails, and is that tested?
Evidence: Error- and exception-path gaps were nearly twice as frequent in AI-co-authored pull requests (470 open-source PRs; authorship inferred from signals) (CodeRabbit, 2025, vendor research). Vendor study with inferred authorship; no peer-reviewed study isolates error handling yet.
FM-07: The unstated requirement
Critical · Moderate evidence
The code does exactly what was asked and not what was needed. The requirement that mattered was never written down, so it was never in the prompt, so it is not in the code. A model can only implement what it is told. People fill unstated requirements from context and experience; a model fills them with the statistically common default.
In review: Nothing looks wrong. The code is clean, the tests pass, and the reviewer approves it against the ticket, which is exactly the problem. It surfaces later as "that's not what I meant".
What catches it:
Discovery before generation, with the decision, the edge cases and the out-of-scope list written down first
Acceptance criteria reviewed by someone who owns the business rule, not only the code
Ask: Which requirement did nobody write down, and how would this code behave if it were violated?
Evidence: "Misinterpretation" (code diverging from the intended behaviour) was the largest category, at 20.77% of 333 buggy samples from CodeGen, PanGu-Coder and Codex; "non-prompted consideration" a further 8.15% (Empirical Software Engineering, 2025). Models struggle to tell an under-specified task from a well-specified one; letting them ask clarifying questions improved performance by up to 74% over non-interactive settings (ICLR, 2026). The research measures misinterpretation inside benchmarks; the business-requirement version is practitioner evidence.
FM-13: Duplication instead of reuse
Material · Moderate evidence
New code re-implements what already exists in the codebase instead of reusing or refactoring it. A model sees the prompt and the open files, not the whole codebase. Generating a fresh copy is always easier than finding the existing function.
In review: A helper that looks familiar because it already exists three directories away, with a slightly different bug.
What catches it:
Duplicate-code detection in CI with a threshold
Reviewer asks "where else does this already exist?"
Ask: Does this already exist in our codebase, and if so, why isn't it being reused?
Evidence: Copy-pasted lines rose from 8.3% to 12.3% of changed lines between 2021 and 2024, while lines associated with refactoring fell from 25% to under 10% (211 million changed lines) (GitClear, 2025, vendor research). Industry-wide trend data, not a direct AI-versus-human comparison.
The code grades itself
Tests written from the implementation describe the code back to itself, so a green pipeline stops being evidence.
FM-10: Tests that bless the bug
Critical · Moderate evidence
Generated tests assert what the code does rather than what it should do, so a bug is encoded as expected behaviour and the suite stays green. When a model writes tests from the implementation, the implementation is its only source of truth. It describes the code back to itself.
In review: High coverage, every test passing, and assertions whose expected values were copied from the output. Nobody can say which requirement a test protects.
What catches it:
Expected values derived from the requirement or a worked example, never from running the code
Mutation testing on critical modules
Tests written (or reviewed) by someone other than the author of the code
Ask: Where did each expected value come from, the requirement or the code's own output?
Evidence: LLM-generated test assertions tend to capture the actual program behaviour rather than the expected behaviour (24 Java repositories) (arXiv preprint, 2024). Preprint; not yet peer-reviewed as far as checked.
FM-11: Passing tests mistaken for quality
Material · Moderate evidence
The team treats "it compiles and the tests pass" as evidence that generated code is fit to ship. Generated code is almost always syntactically correct (Veracode, 2026) and passes the tests generated with it by construction, so the usual signals go green while security and robustness are unmeasured.
In review: A pull request whose only evidence is a green pipeline. The tests check function; nothing checks security, failure handling or the unstated requirement.
What catches it:
Quality gate includes security scanning and failure-path tests, not only unit tests
Definition of done names the evidence required for generated code
Ask: What does this pipeline prove beyond "it runs"?
Evidence: Across five LLMs tested on 4,442 Java tasks, a model's functional correctness (Pass@1) showed no direct correlation with the number of quality and security issues SonarQube found in its passing solutions (arXiv preprint (authors from Sonar), 2025, vendor research). Sabra et al. is a preprint by Sonar authors using Sonar's own tool, and compares five models rather than tasks.
The humans relax
Fluent code reads as competent and arrives faster than it can be traced. Only a review process designed for it catches these.
FM-12: Reviewer over-confidence
Critical · Moderate evidence
People who write code with an AI assistant produce less secure code while believing it is more secure (Perry et al.). No controlled study has yet measured reviewers of generated code; that scrutiny drops on code that looks finished is an inference from that result and from practitioner reports. Fluent, complete-looking code reads as competent. The likely mechanism: trust transfers from the polish of the output to its correctness.
In review: Fast approvals on large generated changes, "LGTM" on code the reviewer did not trace, and no one able to say who checked the risky path.
What catches it:
Named reviewer for generated code, with the risky paths listed in the PR
Review checklist drawn from this register
Size limit on AI-assisted pull requests
Ask: Did the reviewer trace the risky path, or approve the polish?
Evidence: 36% of participants with an AI assistant wrote SQL-injectable code against 7% without (p = 0.041), and AI users were more likely to believe their code was secure (ACM CCS, 2023). 75.8% of surveyed developers and IT managers believed AI-generated code is more secure than human-written code (Snyk, 2023, vendor research). Only 12% of 100 engineering directors and VPs were "very confident" in the quality of AI-generated code before it reaches production, and 62% rely on manual peer review as their primary quality gate (Qodo, 2026, vendor research). Perry et al. is a small 2022-era study; the surveys measure belief, not code quality.
FM-14: The productivity perception gap
Material · Moderate evidence
Teams feel faster with AI assistance and size their review and test capacity to that feeling rather than to measurement. Generating code feels like progress. Reviewing, debugging and integrating it is slower, quieter work that doesn't register as "AI time".
In review: More, larger changes arriving at the same review capacity; review becomes the bottleneck, and the bottleneck gets waved through.
What catches it:
Measure cycle time and escaped defects before and after AI adoption
Review capacity planned for the volume of generated change
Ask: Have we measured the speed-up, or only felt it, and did review capacity grow with the volume?
Evidence: In a randomised trial (16 experienced open-source developers, 246 tasks), early-2025 AI tools made tasks take 19% longer, while developers believed they had been about 20% faster (METR, 2025). METR's February 2026 update says developers are likely more sped up by AI tools now than in early 2025, but that selection effects make its data only very weak evidence for how large that increase is (METR, 2026). Never state "AI makes you slower" as a current fact; the durable finding is the gap between perceived and measured speed.
FM-16: Oversized changes overload review
Material · Practitioner observation
AI-assisted work arrives as fewer, larger changes that exceed what one reviewer can genuinely check. Generation makes large changes cheap to produce; nothing makes them cheaper to review.
In review: A several-thousand-line pull request approved in minutes.
What catches it:
Pull-request size limit with an exception process
Generated changes split by concern before review
Ask: Could one person genuinely review this change in the time it took to approve?
Evidence: No published figure; see the limits. Practitioner observation. Vendor research (Apiiro, reported in the press) describes fewer, larger AI-assisted pull requests, but the primary source could not be opened, so no figure is used.
Score your team against it
The scoring sheet uses the same five levels for every row:
Score | Meaning |
0 | Not considered. Nobody checks for this |
1 | Checked by convention. Some reviewers look; nobody is required to |
2 | Named review question. It's on the checklist and someone owns it |
3 | Automated check. A tool or test detects it on every change |
4 | Gated. The release cannot proceed while it fails |
To use it:
Pick your last AI-assisted change that reached production, not your best one.
Score the control that actually ran on that change, not the one in the policy. If a reviewer would have caught it "if they'd looked", that is a 1.
Read the six critical rows first (FM-01, 02, 07, 08, 10 and 12). Any critical row at 0 or 1 is a gap you are shipping through today.
Fix in mechanism order. Automate the first-mechanism rows (dependency checks, secret scanning, static analysis): cheap, fast and permanent. Then move the second- and third-mechanism rows into acceptance criteria and test design, where mutation testing is the quickest way to find tests that bless the bug. Write review questions only for what's left.
A quick test for the most important row: name the person or the tool that would have stopped FM-07 on your last release. If you can't, it scored 0.
Where the register doesn't apply, and where it isn't enough
Throwaway code. Spikes, prototypes and single-user scripts don't need sixteen checks. The exception ends the moment the code gets a second user, or goes near money, personal data or access control.
AI features. This register is about code a model wrote. If your product calls a model at runtime, you have a different set of failure modes: non-deterministic output, retrieval that quietly returns the wrong context, and reliability that compounds across steps, which is why a 95%-reliable model can build a 60%-reliable system. A companion register for AI features is planned.
People make these mistakes too. The register claims predictability, not inferiority: generated code fails in knowable places, at a volume existing review wasn't sized for. A careful developer given Penhale's one-line ticket might have written the same except branch. What changed is everything around it: the test came from the same source as the bug, the change arrived looking finished, and it was approved at the speed polished code invites.
The controls cost something. Treating deprecation warnings as build failures will block upgrades on noisy libraries. Static-analysis false positives teach people to suppress rules, so tune the rules before you gate on them. Size limits split features across pull requests and lose context, which is why the limit needs an exception process. Mutation testing is slow, so keep it to the modules where a wrong value costs money or trust.
The evidence has limits. Several of the sources are vendor research, named as such. Some studies used models that are now a generation or two old, and one of the strongest (Perry et al.) is small. Each entry says how strong its evidence is. The first-mechanism rates will move as models improve, although Veracode's security pass rate barely moved between its two reports. The other three mechanisms are about how teams specify, test and review, so they age much more slowly.
Making it stick
Give the register an owner, usually the quality lead, and keep it next to the code, not in a wiki. Put the review questions into your pull-request template, but only for the rows you haven't automated; a template with sixteen questions trains people to tick boxes. Set a size limit for AI-assisted pull requests from your own data, not a benchmark: pull your last 20, plot lines changed against review time, and set the limit where review time stops rising with size. Past that point nobody is tracing the code.
Then close the loop. Every escaped defect that traces back to generated code gets mapped to a row. If it fits a row, that row's control failed, and the score drops. If it fits no row, you've found a seventeenth, and the register grows. Review it quarterly against one signal: escaped defects per release that trace back to generated code, compared with the quarter before. If that number is rising, or flat while your scores went up, the scores are flattering you. Quality gates that actually gate covers how to turn the rows you score at 3 into rows you score at 4.
What to do next
Score your last AI-assisted release against the register this week, sitting down with the person who reviewed it.
Download the working checklist and one-page scoring sheet: the full register as a review checklist, the scoring sheet, and an Excel workbook that flags your critical gaps.
If FM-07 scored low, start with discovery for AI features. For the step-by-step review procedure, work through Review AI-Generated Code. For a shorter introduction to why assistants get code wrong, read What LLMs Get Wrong About Code.
Final takeaway
A defensible AI-assisted release is one where, for every critical row, you can name the control that ran. Start with FM-07: if nobody can say what would have stopped it on your last release, that's the first thing to fix.
Sources
Accelerate State of DevOps Report 2024 (DORA). Google Cloud (DORA), 2024, vendor research.
2026 GenAI Code Security Report. Veracode, 2026, vendor research.
Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions (Pearce et al.). IEEE Symposium on Security and Privacy, 2022.
We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs (Spracklen et al.). USENIX Security Symposium, 2025.
LLMs Meet Library Evolution: Evaluating Deprecated API Usage in LLM-based Code Completion (Wang et al.). ICSE, 2025.
Bugs in Large Language Models Generated Code: An Empirical Study (Tambon et al.). Empirical Software Engineering, 2025.
State of AI vs Human Code Generation. CodeRabbit, 2025, vendor research.
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering (Vijayvargiya et al.). ICLR, 2026.
State of Secrets Sprawl 2025. GitGuardian, 2025, vendor research.
Do Users Write More Insecure Code with AI Assistants? (Perry et al.). ACM CCS, 2023.
Do LLMs generate test oracles that capture the actual or the expected program behaviour? (Konstantinou, Degiovanni, Papadakis). arXiv preprint, 2024.
Assessing the Quality and Security of AI-Generated Code: A Quantitative Analysis (Sabra et al.). arXiv preprint (authors from Sonar), 2025, vendor research.
Snyk AI Code Security Report (survey, n = 537), as cited by Snyk. Snyk, 2023, vendor research.
The AI Code Quality Gap (Gatepoint Research survey). Qodo, 2026, vendor research.
AI Copilot Code Quality: 2025 Data Suggests 4x Growth in Code Clones. GitClear, 2025, vendor research.
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. METR, 2025.
METR uplift update. METR, 2026.
LiCoEval: Evaluating LLMs on License Compliance in Code Generation (Xu et al.). ICSE, 2025.


