Coding Agent Ticket Readiness Checklist
Clear tickets and verified codebases prevent AI agents from confidently writing wrong code.

A coding agent does not pause when it hits an ambiguous requirement. It fills the gap and keeps moving, producing code that looks plausible and satisfies neither the intent nor the edge cases its author had in mind. That is not a flaw in the model. The agent is built for autonomous completion, not for stopping to ask a question, and research backs this up directly: models generate code in over 63% of ambiguous scenarios without seeking clarification. The system was trained to keep going, not to stall on a vague instruction, so a vague ticket that a senior engineer would silently complete through years of contextual knowledge becomes, in the agent's hands, a confident, plausible, wrong diff, delivered at machine speed. The error does not surface during generation; it appears later, at review, after the code already exists and someone has to work backward to find out what went wrong and why.
This changes where the real cost of software delivery sits. As AI speeds up the writing of code, implementation time becomes less scarce, while specification clarity and verification capacity become the real constraints. DORA's 2026 research calls this a "verification tax," the cost of checking AI-generated code that eats into the productivity gains AI was supposed to deliver. Individual results diverge across a team. Individual velocity metrics look great, commits pile up, pull requests multiply, but review queues back up behind them, and team-level delivery flattens or gets worse, because every diff now needs a human to reverse-engineer whether the agent actually understood the ticket. An agent is a force multiplier, and what it multiplies is whatever specification it was given. A clear ticket produces fast, correct work. A vague one produces fast, wrong work. The checklist that decides which of those happens lives at the ticket level, and that is the subject of everything that follows.
What "ready" means at the repository level before a single ticket is assigned
No ticket, however well written, can make up for a repository that gives the agent nothing to check its own work against. Ticket quality is the lever that determines whether an agent does the right thing. Repository quality is the floor that determines whether the agent can even tell. Both have to be in place, and neither substitutes for the other.
Larridin's Agent Readiness framework gives a concrete way to measure that floor: 84 binary checks spread across five levels, Baseline, Documented, Agent-ready, Optimized, and Autonomous. The levels are contiguous rather than additive, so a repository has to clear every check at one level before it can credibly claim the next. A codebase with excellent tracing but no committed linter configuration is not "mostly ready": it is still stuck at Baseline, because the foundation underneath the more advanced capability isn't there.
A handful of categories matter most for keeping an agent's output reliable. Style and validation, a committed formatter config, a linter, strict typing, remove entire categories of review churn by narrowing the space of plausible-but-wrong edits before a human ever opens the diff. Testing matters most of all: the 8-pillar verification framework treats it as the foundation, calling for test coverage above 70% on critical paths, a full suite that runs quickly, and a flaky-test rate close to zero. Build systems need to be reproducible, with dependencies pinned and failures that produce clear, actionable error messages, so that an agent's self-check loop terminates on a real signal. Observability, structured logs, error tracking with context, distributed tracing, gives the agent a feedback signal beyond "it compiled." And security gates, static analysis on every pull request, secrets kept out of the codebase, automated and merge-blocking dependency scanning, catch the failure modes that a human reviewer might not notice until it's too late. None of this is a tutorial on how to build a mature engineering org; it is the baseline that every ticket-level instruction below assumes is already in place.
The problem specification: what the ticket must say about what the code should do
Once the repository can give an agent reliable feedback, the first and most common point of failure moves to the ticket itself, specifically to whether it actually says what the code should do. Agents rarely fail because they reason badly or write bad syntax. They fail because they read a requirement and land on a different meaning than the one the author intended, and nothing in the ticket catches the mismatch before code gets written.
A specification that actually works answers four things concretely. It states what the system should do in the affected state, not what it currently does and not what the engineer assumes is obvious. It names the inputs: where they come from, what format they arrive in, and their volume or frequency if that changes the design. It names the output or side effect explicitly, rather than leaving it implied by a two-word task label. And it draws a hard line around the affected component, naming which files or modules are in scope and which are explicitly not.
Consider a ticket that simply says "add pagination." That instruction is satisfied just as completely by an agent that builds infinite scroll as by one that builds explicit page navigation with numbered controls. Both are valid readings of the word "pagination." Only one of them matches what the product actually needs, and the agent has no way to know which one without being told. This is not a hypothetical edge case. Missing file hints, ambiguous scope, and instructions given out of order all measurably raise the odds of an agent misreading what was asked, and these are recurring patterns, not one-off mistakes from a single confused run. A useful test for whether a ticket contains enough to avoid this is whether the team can explain, in writing, how this kind of work happens today, who makes the relevant decisions, and what should happen when something unusual comes up; if they cannot, that information is also missing from the ticket, and the agent will invent an answer in its place.
Acceptance criteria: the contract that tells the agent when to stop
Saying what the code should do answers only half the problem. Telling the agent how it will know the job is finished matters just as much, and that is what acceptance criteria are for. Without them, the agent sets its own definition of "done," and it will either stop short of what was needed or keep going well past it. Both outcomes generate the same downstream cost: more work for whoever reviews the result.
A ticket with no constraints and no acceptance criteria functions as an open invitation to over-build. One real case, drawn from research into agentic coding failures, illustrates the stakes well: when a boundary is never stated, the agent treats its absence as permission to expand the change. Acceptance criteria close that gap only if they're written in a form the agent can actually use. They need to describe verifiable outcomes rather than intentions: "the user sees an error message when the field is empty," not "handle validation." They need to state the negative cases explicitly, meaning what the code must not do, which paths are out of scope, and which edge cases are being deliberately deferred to a later ticket. Each criterion should map to a test case that either the agent or a CI pipeline can run directly; a criterion that only a human can judge by feel is not something an agent can check itself against, no matter how clearly it's written. And the ticket needs a named success state up front, so the agent knows what passing looks like before it starts work, rather than finding out during review that it guessed wrong.
Scope constraint: the boundary that keeps an agent edit from becoming a refactor
Acceptance criteria describe what finished looks like. Scope describes what the agent is allowed to touch to get there, and conflating the two is one of the more common ways a ticket fails. An agent working without an explicit scope boundary tends to make changes that are locally sensible and globally disruptive: it fixes the bug and refactors the surrounding module while it's in there, or updates the function and then adjusts every caller it notices along the way, or adds the requested feature and quietly improves adjacent code nobody asked it to touch.
A scope constraint that actually holds names the files or directories the agent may modify, and separately names the files or directories it must not touch under any circumstance, with CI configuration, authentication, infrastructure, and migration files at the top of that second list. It states what is out of scope for this particular ticket, even when that excluded work looks closely related to what's being asked. And it sets a diff expectation: a reviewer should be able to predict, roughly, what the pull request will contain before opening it. The verification framework turns this into an actual merge gate. It blocks the merge outright when unrequested files, CI configuration, authentication code, or infrastructure files show up in the diff. The scope line drawn in the ticket becomes the first automated check the PR has to pass.
Engineers sometimes resist this kind of constraint on the theory that more context always helps an agent do better work. The evidence points the other way: practitioners have found that an agent's effectiveness actually drops when it's handed too much context to sort through. A precise scope statement isn't a safety rail bolted on to limit the agent. It improves the quality of the signal the agent is working from, which is a different thing from restricting it.
Relevant context the agent cannot infer from the codebase alone
An agent that only reads the codebase will write decisions that are syntactically correct and still inconsistent with choices the team already made, in a meeting, in a Slack thread, in a prior ticket the agent was never shown. Code tells you what the system does now. It does not tell you why it got that way, what was tried and abandoned, or who needs to approve a change before it ships.
A handful of categories of information belong in a ticket precisely because they live outside the repository and nowhere else. The reason the change is being made shapes the right solution more than the task label or the surface request does. Prior attempts matter too: if an approach was tried before and reverted, the ticket has to say so and explain why, since an agent with no memory of that history will walk straight back into the same dead end. Constraints invisible in the code itself need to be spelled out: a performance requirement nobody wrote down, a third-party API's contractual limits, a regulatory restriction, a design decision that was made deliberately against the more obvious implementation. And the ticket should name the relevant decision-makers, who approved the approach being asked for, and whose sign-off the resulting pull request will actually need before it can merge.
Some of this belongs in a persistent reference used across every ticket. The AGENTS.md pattern is built for exactly that: conventions, project state, and known past bugs that apply across all tickets, written once and reused. But context specific to a single piece of work, why this change, why now, what was already tried and failed, cannot be encoded once and pulled in automatically. It has to live in the ticket itself. The verification framework's documentation pillar states this principle at the level of code comments: they should explain why, not just what. The same logic applies to the ticket one level up. The agent needs the reasoning behind the request, not only the instruction itself.
The testing and verification gate: what the ticket must specify about proof
A ticket that doesn't say how its output will be verified hands that decision to the same system that wrote the code in the first place. That creates a self-grading problem: the agent produces the implementation and, left unchecked, also produces the evidence that the implementation is correct. A proof written by the party being tested is not proof.
Closing that gap means the ticket has to state explicitly what kind of verification is required. It should name which test types apply, unit, integration, contract, end-to-end. It should state which existing tests must keep passing and flag any that are expected to change as part of this work. It should say whether new tests are required and specify what behavior they need to cover as a firm requirement. And it should name who owns the verification sign-off, a specific person, not "CI" standing in as a substitute for human judgment.
These requirements feed into a set of merge gates that give the review process actual teeth. A merge should be blocked when the diff doesn't match the ticket and its approved files, when a deterministic check fails, when a stated requirement or negative case fails, when tests were deleted or quietly weakened, when a high-risk security finding or an exposed secret remains in the diff, when a new or unverified dependency shows up uninvited, or when there's no named owner, no rollback path, or no way to trace the change back to its ticket. None of these gates work if the ticket never specified what proof looked like in the first place. The gate only functions because the ticket gave it something to check against.
Authority boundaries: changes that require a human sign-off regardless of ticket quality
Ticket completeness and agent authority are two separate questions, and a ticket can score perfectly against everything above and still describe a change that should not go forward without a human making the final call. A well-specified ticket tells the agent what to build. It does not, by itself, settle whether the agent should be the one building it unsupervised.
The agent authority checkpoint frames this as a set of distinct capabilities: what the agent is allowed to read, what it may suggest, what it may prepare, what it may execute outright, and what it must never touch no matter how clear the instructions are. A ticket needs to declare which of those categories the work falls into before any code gets written. Certain categories carry enough downside that they call for a human in the loop regardless of how airtight the specification is: authentication and authorization logic, payment processing and any path touching financial data, infrastructure and deployment configuration, database migrations with destructive potential, and public API contracts, whether the change is an addition, a modification, or a deprecation. Personal data handling and anything touching compliance scope belong on this list too, and the reason is measurable: AI-generated code introduces vulnerabilities in 45% of tested tasks, which makes independent human review of compliance-adjacent work a requirement rather than a precaution.
The underlying security principle is a staged permission model. The agent plans with read-only access, implements with write access confined to its approved scope, verifies through tests run in a sandbox, and a separate human or CI identity, independent of the agent, handles the actual release and deployment. That separation of stages is what stops an agent from writing, testing, and approving its own high-risk change as one uninterrupted action with no outside check at any point. A complete ticket tells the agent what correct looks like. The authority boundary decides whether the agent gets to act on that answer alone, and for the categories above, the answer stays no, no matter how good the ticket is.

Sources
- AI Agent Readiness Checklist — 84 Repository Checks
- AI Coding Agent Readiness Checklist · GitHub
- How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions
- 2026 Agentic Coding Trends - Implementation Guide (Technical)
- GitHub - ti-kiu/agent-readiness-checklist: Open-source checklist and implementation guide for achieving Agent-Native status (Level 5) on isitagentready.com. Covers robots.txt, MCP, OAuth, DNS-AID, WebMCP, and more. · GitHub
- feat(aidd-dev): assess repository readiness for coding agents · Issue #954 · ai-driven-dev/framework
- Ticket-authoring conventions: scope declaration block, wave and gate labels, claim and escalation comment formats, model recommendation line · Issue #720 · socket-link/ampere
- triage-labels: a ticket's premises hold before it gets ready-for-agent · Issue #360 · corygyarmathy/dotfiles
