What AI-Generated Code Structurally Skips
Generated code builds what the ticket says. The failures live in what the ticket never said, and they cluster in the same handful of places.
The ticket read: "Add an endpoint so support agents can refund an order." An engineer gave it to a coding assistant, which produced a tidy FastAPI route and two tests. Both tests passed. Review took four minutes because there was almost nothing to object to: clear names, a 404 for a missing order, a commit at the end. It merged on a Tuesday.
On Thursday, someone in finance noticed a refund issued twice on the same order, eleven seconds apart. On Friday, a security reviewer noticed that the endpoint checked whether the caller was logged in but never whether they were a support agent, so any authenticated customer could refund any order whose ID they could guess.
(Composite scenario, assembled from failure patterns common in AI-assisted delivery. Not a single named incident.)
Nothing in that code was wrong in the way a linter or a syntax check understands "wrong." Every line did what it said. The defects were all absences: a check never written, a case never handled, a record never kept. That's the shape of AI-generated code failure, and once you can see the shape you can gate for it.
Better review is the obvious answer, and it doesn't hold
The instinct after an incident like that is "review AI code more carefully." It's reasonable. Reviewers catch things, and line-by-line review of generated code is a real skill (we teach it step by step in Review AI-Generated Code).
It doesn't scale as the gate, for a structural reason. Review is good at spotting what's there and wrong, and bad at spotting what's missing. Nothing in a clean ten-line function prompts "where's the role check?" The absence has no line number.
The evidence lines up with that. CodeRabbit's analysis of 470 open-source pull requests (December 2025) found AI-co-authored PRs carried about 1.7× more issues than human-only ones: 10.83 per PR against 6.45. Error-handling and exception-path gaps were nearly 2× more common. Security issues ran up to 2.74× higher, and the most frequent were improper password handling and insecure object references, which is exactly the refund bug above. CodeRabbit is a code-review vendor and inferred authorship from signals rather than confirming it, so treat the ratios as directional. The categories are the useful part: they're omissions, not typos.
Test you can run: pull your last five merged AI-assisted PRs and, for each one, write down the review comments that named something missing versus something wrong. If the "missing" column is near zero, your review process isn't seeing absences. That doesn't mean there weren't any.
The model builds the spec, and the spec is incomplete
Here's the mechanism. A coding model completes the task it was given. The ticket states a happy path: agents refund orders. The model implements that path well. It does not know, because nobody told it:
who counts as a support agent, and whether agents are scoped to a tenant or region
that the payment provider can time out after moving money
that browsers double-submit and clients retry on a 504
that two agents might open the same escalation at once
that finance reconciles refunds against an audit trail
Each of those is a requirement, and none of them is in the ticket. A senior engineer fills them in from domain knowledge without noticing. The model fills them in with nothing.
Humans skip these too. Every category below also appears in hand-written incident reports. The difference is rate and visibility. A human who forgot the role check usually forgot it on a bad day. A model forgets it whenever the prompt doesn't mention it, and it produces code clean enough that nobody notices.
Then the model writes tests from the same prompt, so they cover the same stated path: refund succeeds, missing order returns 404. Code and tests share one set of blind spots. A green suite here proves the code matches the ticket. It says nothing about whether the ticket matched reality.
That's why this is a discovery problem, not a tooling one. You can't test your way to a requirement you never captured, and you can't generate your way to one either. The model is a fast, literal implementer of whatever discovery produced. When discovery produced one sentence, you get one sentence of behavior.
The same holds one level up. Architecture decisions nobody makes on purpose still get made, by whoever writes the first code, which is the argument of Why Architecture Matters Before a Single Line of Code. With a coding model in the loop, that first code arrives in minutes.
There is research support for the "less than it should be" shape. A September 2026 preprint by Improta, Liguori and Cotroneo compared 787,562 human-written and AI-generated function pairs across Python, Java and C. The AI implementations were roughly half the size and half the branching of the human ones. Branching is where code handles cases. Half the branches doesn't automatically mean half the cases handled, since some human branching is accumulated cruft. But it fits what reviewers keep finding: generated code takes the main road and skips the side streets.

Test you can run: take one AI-authored function and list every requirement it satisfies. Then mark each one either stated in the ticket or prompt or inferred by the model. If nearly everything is in the first column, the model only built what it was told, and anything you didn't tell it is missing.
Why better models haven't closed the gap
If this were a capability problem, newer models would be closing it. Veracode's data says they aren't. Its GenAI Code Security research gives 100+ models the same 80 coding tasks in Java, Python, C# and JavaScript, with no security-specific prompting. That's the condition that matters here: the security requirement is left unstated, exactly as it is in most tickets. The 2025 report found 45% of samples introduced an OWASP Top 10 flaw. The 2026 report (July 2026) puts the security pass rate at 56%, a figure it describes as stalled.
The divergence is the telling part. Veracode's Spring 2026 update tracked syntax pass rates climbing from about 50% to 95% since 2023, while security pass rates stayed flat between 45% and 55%. By July, syntactic correctness was near 100%. Models got dramatically better at the part of the job the prompt specifies, code that compiles and does the named thing, and barely moved on the part it doesn't. The July report also found no correlation between model size and security. The best model scored 68%, which still means roughly one task in three shipping a known flaw.
The per-weakness numbers point the same way (Spring 2026). SQL injection had an 82% pass rate and insecure cryptographic algorithms 86%. Those are well-known patterns with a canonical safe form. Cross-site scripting passed only 15% of the time and log injection 13%, because defending against them depends on context: where this value ends up, who reads that log. Context is exactly what the ticket didn't carry.
Interpretation, not established fact: the gap that remains is the gap between what's written down and what's known. Model upgrades narrow the first. Only better intake narrows the second.
Test you can run: check which model version produced your last month of AI-assisted code. If your gate assumes "the new model is safer," find the evidence you're relying on. If it's a launch blog post, that's a vendor claim, not a measurement of your codebase.
Asking the model to fix its own code makes it worse
The second instinct is to put the model in the loop: "review this for security issues and fix them." Research suggests that can make things worse. Shukla, Joshi and Syed (arXiv, 2025) ran 400 code samples through 40 rounds of LLM "improvement" and found a 37.6% increase in critical vulnerabilities after five iterations. Without an external oracle, iteration drifts.
AI review still has a place as a second reader with a checklist. It fails as the gate because it shares the original model's missing context. If the ticket didn't say refunds are tenant-scoped, the reviewing model doesn't know that either.
Test you can run: if your pipeline has an "AI fixes AI findings" loop, count how many rounds a typical PR goes through and diff the first and last versions. If lines changed keeps growing past round two, you're paying for drift.
Walking one PR through the gate
Here is the refund endpoint in full, as generated. (Composite code walkthrough, written to show the pattern. It is representative of real assistant output, but it isn't a captured transcript.)
@router.post("/orders/{order_id}/refund")
def refund_order(order_id: int, body: RefundRequest,
user: User = Depends(get_current_user),
db: Session = Depends(get_db)):
order = db.get(Order, order_id)
if order is None:
raise HTTPException(status_code=404, detail="Order not found")
payments.refund(order.payment_id, body.amount)
order.status = "refunded"
order.refunded_amount = body.amount
db.commit()
return {"status": "refunded", "amount": body.amount}
class RefundRequest(BaseModel):
amount: float
The tests it wrote were test_refund_success and test_refund_order_not_found. Both passed.
Now put it through the seven categories. For each one, the reviewer must point at the line or test that handles it, or write down why it doesn't apply. "Looks fine" doesn't count.
Who may call it? get_current_user authenticates. Nothing authorizes. No role check, no tenant check. Any logged-in customer can refund order 4471 by changing a number in the URL. This is an insecure direct object reference, the category CodeRabbit flagged as most common among AI security issues. Absent.
What if the input is hostile? amount: float. A negative amount, an amount larger than the order total, and floating-point money are all accepted without complaint. Absent.
What if it fails halfway? The provider call happens before the database commit. If the commit fails, the money has left and the order still reads paid, so support will refund again. If the provider times out, nobody knows whether the money moved. Absent.
What if it runs twice? No idempotency key. A double-click or a client retry on a gateway timeout issues two refunds. This was Thursday's incident. Absent.
What if it runs at the same time? Two agents on the same escalation both read status = paid and both refund. No row lock, no state check. Absent.
What gets recorded? No audit event: no record of who refunded, how much, why, or the provider's reference. Finance will ask, and the answer will be "check the application logs." Absent.
Limits and cost? One provider call per request, no loops, no unbounded queries. Not applicable, and that's stated. The gate doesn't have to find seven problems every time. It has to account for seven categories every time.
Six of seven absent, in code that passed its tests and its review. What closes the gap is tests the model didn't write, aimed at the categories it skipped:
def test_customer_cannot_refund_another_customers_order(client, customer_token, other_order):
r = client.post(f"/orders/{other_order.id}/refund",
json={"amount": "10.00"}, headers=auth(customer_token))
assert r.status_code == 403
def test_refund_cannot_exceed_refundable_balance(client, agent_token, order):
r = client.post(f"/orders/{order.id}/refund",
json={"amount": str(order.captured + 1)}, headers=auth(agent_token))
assert r.status_code == 422
def test_same_idempotency_key_refunds_once(client, agent_token, order, fake_payments):
headers = {**auth(agent_token), "Idempotency-Key": "esc-8812"}
client.post(f"/orders/{order.id}/refund", json={"amount": "10.00"}, headers=headers)
client.post(f"/orders/{order.id}/refund", json={"amount": "10.00"}, headers=headers)
assert fake_payments.refund_calls == 1
All three fail against the generated code. That's the point: they turn absences into red tests, which reviewers can see.
The cheaper fix happens a step earlier. Here is the block that, added to the ticket before generation, would have produced different code. We call it the unstated-requirements block:
Who may do this: role=support_agent, orders in the agent's own tenant only. Everyone else: 403.
Hostile input: amount is Decimal, > 0, <= captured minus already refunded. Else 422.
If it fails halfway: write refund as PENDING before the provider call; reconcile on timeout.
If it runs twice: require Idempotency-Key; same key -> same response, one provider call.
If it runs at once: lock the order row; second concurrent request gets 409.
What gets recorded: audit event {actor, order_id, amount, reason, provider_ref, timestamp}.
Limits/cost: n/a (one provider call).
Seven lines, about ten minutes with someone who knows the domain. Most payment APIs accept an idempotency key (Stripe's does, for example), so line four is usually configuration, not invention. Given this block, a current coding model will generally write the role check, the lock and the pending state, because now they're part of the task.
The gate, category by category
This is the working tool. Use the left two columns to train reviewers, the third as the PR checklist, and the fourth as the ticket template.
What gets skipped | Why the model skips it | The check on the PR | The spec line that prevents it |
Authorization (who may do this, to which records) | Authentication is visible in the framework. Business roles and tenancy live in people's heads | A test where the wrong role and the wrong tenant each get 403 | "Who may do this: role, scope, everyone else gets…" |
Validation at trust boundaries | The happy-path example uses valid input. Nothing signals hostility | Tests for negative, oversized, malformed and wrong-type input at every external entry point | "Hostile input: type, bounds, rejection code" |
Partial failure | Side effects look atomic in code, and the prompt never mentions timeouts | Identify every external call. Show what state persists if it fails before or after | "If it fails halfway: what state, how reconciled" |
Idempotency | Retries and double-submits come from the environment, not the feature | A test that sends the same request twice and asserts one effect | "If it runs twice: key, response, effect count" |
Concurrency | Tests run serially, so races never appear in generated tests | Identify shared mutable state. Show the lock, version check or constraint | "If it runs at once: lock or conflict behavior" |
Audit and observability | Nobody asks for logs in a feature ticket, but someone needs them later | Show the audit event or structured log, with actor and correlation ID | "What gets recorded: event and fields" |
Resource limits and cost | Unbounded loops and N+1 queries work fine on test data (CodeRabbit found excessive I/O ~8× more often in AI PRs) | Check for pagination, timeouts and query counts on realistic volumes | "Limits: max rows, timeout, rate" |
A note on scanners. Static analysis catches pattern-shaped flaws well. That's how Veracode measures SQL injection and XSS. It is weak on the first, third, fourth and fifth rows, because nothing in the code is syntactically wrong. A missing authorization check is invisible to a tool looking for bad patterns. Keep your scanner, and don't let it stand in for this table.
Test you can run (the skip audit): take your last ten AI-assisted PRs. For each, score the seven rows present, absent, or n/a with a reason. Count the "absent" cells. Our rule of thumb, not a benchmark: more than about ten absences across the 70 cells means the gate is needed now. If the absences cluster in one row, you've found your first template line.
When a lighter gate is fine, and when it isn't
The full gate costs time: roughly ten minutes of intake per ticket, plus a reviewer accounting for seven categories. That isn't free, and it isn't always warranted.
Context | Gate | Why | Warning sign it's the wrong call |
Spike or throwaway prototype | None or rows 1–2 only | The code is meant to be deleted | The prototype gets a production URL "temporarily" |
Internal script, no trust boundary, single operator | Rows 3, 6, 7 | No hostile input or multi-user access, but failures and cost still matter | It gets scheduled, shared or pointed at production data |
Repo with strong existing middleware (central authz, idempotency layer) | Full gate, but rows 1 and 4 may be one-line confirmations | The model may correctly reuse existing decorators | The skip audit still shows absences in those rows. Check, don't assume |
Customer-facing, money, PII, regulated data | Full gate plus independent tests | Absences here become incidents and findings | — |
The prototype row deserves extra attention. Vibe coding is a legitimate way to explore. The failure mode is the prototype that quietly becomes the product with none of the seven categories ever revisited.
Test you can run: list every repo where AI-assisted code merged in the last quarter and put each one in a row of the table above. Any repo you can't place is a repo nobody has decided the risk level for.
Installing the gate: five steps
Label AI-authored work. Use a PR label or a commit trailer (Co-authored-by: is common). Georgia Tech's Vibe Security Radar traced 74 public CVEs to AI coding tools by March 2026, with 6, 15 and 35 disclosed in January, February and March. It can only see commits carrying an AI signature, and its researchers estimate the true count is five to ten times higher. Your organization has the same blind spot: you can't measure the skip rate of code you can't identify. Owner: tech lead. Done when ≥90% of AI-assisted PRs carry the label for two weeks running.
Add the unstated-requirements block to the ticket template. Seven lines, filled in by the ticket author with whoever knows the domain. "n/a" is allowed only with a reason. This sits naturally alongside your acceptance criteria. It's the same discipline, aimed at what acceptance criteria usually leave out. Owner: product owner or BA, with the QA lead. Done when the template field exists and is required.
Feed the block to the model. Paste it into the prompt, and put the standing defaults (your authz pattern, idempotency convention, audit event shape) in the repo's agent instructions file, whatever your tool calls it. Owner: tech lead. Done when the defaults file exists and is reviewed like code.
Make the seven rows a required PR checklist with evidence. Each row needs a link to the line or test, or a reason for n/a. A tick with no link doesn't count. Owner: QA lead. Done when merge is blocked on an incomplete checklist.
Require at least one independent test per "present" row on high-risk changes. Independent means it isn't written from the same prompt that generated the code: a human, or a separate model session working from the requirements block rather than the code. Owner: QA lead. Done when high-risk PRs show tests authored separately from the implementation.
The work moves upstream. Ticket authors will push back on the extra ten minutes because "the AI handles that." The skip audit is your answer. In return, "is AI code safe here?" stops being a debate about models and becomes a count: absences per PR, by category, trending down. DORA's 2025 research found AI adoption now tracks with higher delivery throughput and higher delivery instability, with more change failures and more rework. Teams sped up before their systems could absorb the speed. This gate is one of the cheapest pieces of that missing system.
How the gate itself fails
Checklist theater. Rows ticked with no linked evidence. Detect: sample five merged PRs a month and follow every link. Prevent: no link, no merge.
Mirror tests. The model writes tests for the checklist that assert whatever the code does, bugs included. Detect: a "present" row whose test was generated in the same session as the code. Prevent: step 5.
Reflexive n/a. The requirements block fills up with "n/a" to save time. Detect: more than two n/a rows on a customer-facing ticket. Prevent: n/a requires a one-line reason a reviewer can dispute.
Unlabelled bypass. People stop labelling AI work to avoid the gate. Detect: AI-label rate drops while tool usage (from licence telemetry) doesn't. Prevent: keep the gate lightweight enough that labelling isn't punished. That's what the tiered table is for.
Final takeaway
AI-generated code fails in a predictable shape. It is complete with respect to what you said and silent about what you didn't. Better models have made the stated part nearly flawless and left the unstated part roughly where it was. That makes this a discovery problem with a discovery fix: say the seven things before generation, and account for the seven things before merge.
Your next action: this week, run the skip audit on your last ten AI-assisted PRs: seven rows each, present, absent or n/a with a reason. The row with the most absences becomes the first line of your ticket template. (Choosing a tool rather than gating one? Our honest comparison of Cursor, Copilot and Claude Code covers how they handle repo context. The gate applies to all of them.)
Sources
Veracode, 2025 GenAI Code Security Report (July 2025): 80 tasks, 100+ LLMs, Java/Python/C#/JavaScript; 45% of samples introduced OWASP Top 10 flaws. https://www.veracode.com/resources/genai-code-security-report/
Veracode, Spring 2026 GenAI Code Security Update: Despite Claims, AI Models Are Still Failing Security (24 March 2026): syntax vs security pass-rate trend, per-CWE rates. https://www.veracode.com/blog/spring-2026-genai-code-security
Veracode, 2026 GenAI Code Security Report press release (28 July 2026): 56% security pass rate, no security-specific prompting, no size–security correlation, best model 68%. https://www.veracode.com/news/llms-are-getting-smarter-but-not-safer-veracode-2026-genai-code-security-report-finds-ai-generated-code-security-has-stalled-at-56%25-pass-rate/
CodeRabbit, State of AI vs Human Code Generation Report (17 December 2025): 470 PRs; authorship inferred. Vendor research. https://coderabbit.ai/blog/state-of-ai-vs-human-code-generation-report
Improta, C., Liguori, P., Cotroneo, D., What is the Difference Between Me and You? Benchmarking the Quality Gap Between Human-Written and AI-Generated Code, arXiv:2609.12708 (September 2026, preprint). https://arxiv.org/abs/2609.12708
Shukla, S., Joshi, H., Syed, R., Security Degradation in Iterative AI Code Generation — A Systematic Analysis of the Paradox, arXiv:2506.11022 (2025). https://arxiv.org/abs/2506.11022
Georgia Tech SSLab, Vibe Security Radar, as reported in Security Researchers Sound the Alarm on Vulnerabilities in AI-Generated Code, Infosecurity Magazine (2026). https://www.infosecurity-magazine.com/news/ai-generated-code-vulnerabilities/
DORA / Google Cloud, 2025 State of AI-assisted Software Development (September 2025). https://cloud.google.com/devops/state-of-devops


