AI Engineering
"The Checkout Page Is Not Working": What a Vague Prompt Cost Me
October 10, 2026
The checkout page in a Next.js application I work on was broken: a single click on the checkout button ran the checkout twice. I could see it on the screen. So I did what felt natural and typed this into my coding agent:
The checkout page is not working. Detect and fix it.No file. No mention of the double submission. No idea of what "fixed" should look like. I was vibe coding: I handed over the symptom in the vaguest possible words and expected the agent to find the rest.
In this post:
- What that one-line prompt actually cost
- Why the agent behaved the way it did, and why it wasn't the model's fault
- The framework I found for briefing an agent, and why you should rarely use all of it
- The second attempt, and what changed
What the vague prompt cost
The agent did not ask a single question. It started reading files, built its own theory of what "not working" meant, and acted on it. The result was bad in four ways:
- It burned tokens. The agent searched widely through the codebase because nothing told it where to look.
- The bug was still there. It fixed a problem, just not the one I was looking at.
- It caused side effects. The change spread across multiple files, most of which had nothing to do with the bug, and other parts of the app started behaving differently.
- The code quality was poor. What it did change was forced into working rather than fixed properly: patches around the symptom instead of a repair of the cause.
I had spent tokens and review time, and I was now further from a fix than when I started, because I had to undo the side effects first.
Why it happened
It's tempting to blame the model. That would be the wrong lesson.
Every piece of information missing from a prompt is a decision the agent has to make on its own. My prompt left almost every decision open, so the agent filled each gap with a plausible guess. That is exactly what a language model is built to do. Here is what my one line left out, and what the agent guessed instead:
| What the prompt didn't say | What the agent had to guess | What it led to |
|---|---|---|
| Where the problem is | Which of many files to read and change | Token burn, edits in the wrong place |
| What "not working" means | Which symptom to fix, with no hint of the double submission | A fix for a different problem |
| What must not change | Nothing was off-limits | Side effects in unrelated code |
| What a good fix looks like | Any change that seems to work is acceptable | Patches instead of a real fix |
| How to check the result | Nothing to verify against | No way for the agent to notice it failed |
None of these is about intelligence. A senior engineer given the same sentence with no access to my screen would have asked five questions before touching anything. The agent didn't ask because I never told it that asking was an option.
Finding a better way to brief an agent
After that run, I went looking for a more disciplined way to write prompts for agents. The most useful thing I found was a context-engineering framework that splits a good brief into ten elements, grouped by the life of a task:
| Phase | Element | The gap it closes |
|---|---|---|
| Framing | 1. Agent role | The level of rigor to apply |
| 2. First-line insight | Why the task matters, in one sentence | |
| Grounding | 3. References | Where the truth lives: file paths, tickets, designs |
| 4. Ask-user gate | What to do when something is still unclear: ask | |
| Sizing | 5. Scope signal | Just act, or plan first |
| Execution | 6. Acting | The task itself, as concrete steps |
| 7. Avoid Zone | What must not be touched | |
| Closure | 8. Result format | The shape of the deliverable |
| 9. Definition of Done | How the agent checks its own work | |
| Fallback | 10. Conflict priority | Which instruction wins when two disagree |
Line this up against the table in the previous section and every failure from my first attempt has a matching element. Missing location is References (3). Unclear symptom is Acting (6). Side effects are the Avoid Zone (7). Poor quality and no verification are the Definition of Done (9).
The part that surprised me: don't use all ten
My first instinct was to turn this into a template and fill in all ten fields for every prompt. The framework argues against exactly that, and I think it's the most important idea in it: the ten elements are a ceiling, not a floor.
The test for including an element is not "is it on the list." It is: does leaving this out create a real risk of the agent going the wrong way? Applied honestly:
- Every task needs References (3) and Acting (6). Location and instruction never go away.
- Unclear requirements add the ask-user gate (4).
- Code that could break other things adds the Avoid Zone (7) and a Definition of Done (9).
- Large or risky changes add a scope signal (5) and a conflict priority (10).
- Judgment calls add a role (1) and a first-line insight (2).
A one-line rename needs only a file path and an instruction. A ten-part brief for that rename would bury the instruction and waste the reader's attention, whether the reader is a person or a model. My checkout bug sat in the middle: it needed location, a precise symptom, a boundary, and a way to verify, but not a migration plan.
The second attempt
With that in mind, I rewrote the prompt. Here it is, with each part tagged by the element it covers:
You are a senior frontend engineer. [1]
Read the checkout page, src/checkout/page.tsx, at the checkout
function. [3]
Check how this function handles being triggered twice when the user
clicks once, and fix it. [6]
Include an explanation of the fix for me. [8]
Follow the project's coding conventions. [9]
AVOID side effects, new approaches, and TypeScript errors. [7, 9]Each part closes one gap from the first attempt:
- The file path and the function name tell the agent exactly where to look, so it doesn't search the whole codebase. This is the single biggest change.
- The precise symptom, one click running checkout twice, replaces "not working." The agent now knows which problem to solve instead of inventing one.
- "Avoid side effects and new approaches" is the Avoid Zone. It tells the agent to fix the bug inside the existing design instead of redesigning the page around it.
- "Follow the coding conventions" and "no TypeScript errors" set a quality bar, which is the part of the Definition of Done the first attempt had nothing to measure against.
- "Include an explanation" is the result format. It turns the change into something I can review and learn from, instead of a diff I have to reverse-engineer.
Six elements out of ten. There's no scope signal, no conflict priority, and no first-line insight, because a bug in one function carries none of those risks.
Looking back, I would add two more sentences. The prompt has no ask-user gate, so if the cause had been somewhere other than the checkout function, the agent would have had to guess again. And the quality bar says nothing about behavior: "verify that one click submits exactly once" would have given the agent a concrete way to prove the fix works. It worked without them because the location and symptom were precise, but precision is what made the gap small, not the absence of the gap.
What changed
The difference was immediate:
| First attempt | Second attempt | |
|---|---|---|
| Prompt | "The checkout page is not working. Detect and fix it." | File, function, exact symptom, boundaries, quality bar |
| Files changed | Multiple, most of them unnecessary | One: src/checkout/page.tsx |
| The bug | Still there | Fixed |
| Side effects | Yes, in unrelated parts of the app | None |
| What I got back | A diff to untangle | A focused diff plus a report explaining the fix |
The prompt was longer, but the run was shorter. A few extra sentences up front replaced a wide search, a wrong fix, and a cleanup.
The bigger change was in how I work. I no longer treat a prompt as a wish. I treat it as a brief to a capable colleague who can't see my screen: what I saw, where it lives, what not to touch, and how we'll both know it's fixed.
Where the framework falls short
The framework helped, but it isn't complete, and it's worth knowing where it bends.
- "Stop and ask" assumes someone is watching. For a long unattended run, nobody answers the question. A better instruction there is a safe default plus a visible flag: "leave it unchanged, add a TODO, and list it in the summary."
- The role line does less than people think. "You're a frontend engineer" is the element people write first and the one that does the least work. The file paths and the symptom carry the weight. If you're dropping one element, drop this one.
- Project instructions should carry the permanent rules. If your agent reads a
CLAUDE.mdorAGENTS.md, put your conventions and permanent boundaries there. A per-task prompt should hold only the facts of this task. - It covers what to add, not what to leave out. Pasting a whole log or a long chat history competes with the instruction for attention. Removing noise is part of context engineering too.
Summary
- A vague prompt doesn't save time. It moves every decision to the agent, which fills each gap with a guess, and you pay for the guesses in tokens, side effects, and rework.
- Most bad agent output is a briefing problem, not a model problem.
- A good brief has up to ten elements, but almost no task needs all ten. Always include where and what, then add one element for each risk you can name.
- For bug fixes, the elements that matter most are the exact symptom, the location, a boundary around code that must not change, and a way to verify the fix.
The shortest version fits on one line: say where, say what, say how to check it, then add one sentence for each risk you can name.