Skip to main content
Arthur Ha

AI Engineering

"The Checkout Page Is Not Working": What a Vague Prompt Cost Me

October 10, 2026

The checkout page in a Next.js application I work on was broken: a single click on the checkout button ran the checkout twice. I could see it on the screen. So I did what felt natural and typed this into my coding agent:

Plain text
The checkout page is not working. Detect and fix it.

No file. No mention of the double submission. No idea of what "fixed" should look like. I was vibe coding: I handed over the symptom in the vaguest possible words and expected the agent to find the rest.

In this post:

  • What that one-line prompt actually cost
  • Why the agent behaved the way it did, and why it wasn't the model's fault
  • The framework I found for briefing an agent, and why you should rarely use all of it
  • The second attempt, and what changed

What the vague prompt cost

The agent did not ask a single question. It started reading files, built its own theory of what "not working" meant, and acted on it. The result was bad in four ways:

  • It burned tokens. The agent searched widely through the codebase because nothing told it where to look.
  • The bug was still there. It fixed a problem, just not the one I was looking at.
  • It caused side effects. The change spread across multiple files, most of which had nothing to do with the bug, and other parts of the app started behaving differently.
  • The code quality was poor. What it did change was forced into working rather than fixed properly: patches around the symptom instead of a repair of the cause.

I had spent tokens and review time, and I was now further from a fix than when I started, because I had to undo the side effects first.

Why it happened

It's tempting to blame the model. That would be the wrong lesson.

Every piece of information missing from a prompt is a decision the agent has to make on its own. My prompt left almost every decision open, so the agent filled each gap with a plausible guess. That is exactly what a language model is built to do. Here is what my one line left out, and what the agent guessed instead:

What the prompt didn't sayWhat the agent had to guessWhat it led to
Where the problem isWhich of many files to read and changeToken burn, edits in the wrong place
What "not working" meansWhich symptom to fix, with no hint of the double submissionA fix for a different problem
What must not changeNothing was off-limitsSide effects in unrelated code
What a good fix looks likeAny change that seems to work is acceptablePatches instead of a real fix
How to check the resultNothing to verify againstNo way for the agent to notice it failed

None of these is about intelligence. A senior engineer given the same sentence with no access to my screen would have asked five questions before touching anything. The agent didn't ask because I never told it that asking was an option.

Finding a better way to brief an agent

After that run, I went looking for a more disciplined way to write prompts for agents. The most useful thing I found was a context-engineering framework that splits a good brief into ten elements, grouped by the life of a task:

PhaseElementThe gap it closes
Framing1. Agent roleThe level of rigor to apply
2. First-line insightWhy the task matters, in one sentence
Grounding3. ReferencesWhere the truth lives: file paths, tickets, designs
4. Ask-user gateWhat to do when something is still unclear: ask
Sizing5. Scope signalJust act, or plan first
Execution6. ActingThe task itself, as concrete steps
7. Avoid ZoneWhat must not be touched
Closure8. Result formatThe shape of the deliverable
9. Definition of DoneHow the agent checks its own work
Fallback10. Conflict priorityWhich instruction wins when two disagree

Line this up against the table in the previous section and every failure from my first attempt has a matching element. Missing location is References (3). Unclear symptom is Acting (6). Side effects are the Avoid Zone (7). Poor quality and no verification are the Definition of Done (9).

The part that surprised me: don't use all ten

My first instinct was to turn this into a template and fill in all ten fields for every prompt. The framework argues against exactly that, and I think it's the most important idea in it: the ten elements are a ceiling, not a floor.

The test for including an element is not "is it on the list." It is: does leaving this out create a real risk of the agent going the wrong way? Applied honestly:

  • Every task needs References (3) and Acting (6). Location and instruction never go away.
  • Unclear requirements add the ask-user gate (4).
  • Code that could break other things adds the Avoid Zone (7) and a Definition of Done (9).
  • Large or risky changes add a scope signal (5) and a conflict priority (10).
  • Judgment calls add a role (1) and a first-line insight (2).

A one-line rename needs only a file path and an instruction. A ten-part brief for that rename would bury the instruction and waste the reader's attention, whether the reader is a person or a model. My checkout bug sat in the middle: it needed location, a precise symptom, a boundary, and a way to verify, but not a migration plan.

The second attempt

With that in mind, I rewrote the prompt. Here it is, with each part tagged by the element it covers:

Plain text
You are a senior frontend engineer.                                     [1]
Read the checkout page, src/checkout/page.tsx, at the checkout
function.                                                               [3]
Check how this function handles being triggered twice when the user
clicks once, and fix it.                                                [6]
Include an explanation of the fix for me.                               [8]
Follow the project's coding conventions.                                [9]
AVOID side effects, new approaches, and TypeScript errors.              [7, 9]

Each part closes one gap from the first attempt:

  • The file path and the function name tell the agent exactly where to look, so it doesn't search the whole codebase. This is the single biggest change.
  • The precise symptom, one click running checkout twice, replaces "not working." The agent now knows which problem to solve instead of inventing one.
  • "Avoid side effects and new approaches" is the Avoid Zone. It tells the agent to fix the bug inside the existing design instead of redesigning the page around it.
  • "Follow the coding conventions" and "no TypeScript errors" set a quality bar, which is the part of the Definition of Done the first attempt had nothing to measure against.
  • "Include an explanation" is the result format. It turns the change into something I can review and learn from, instead of a diff I have to reverse-engineer.

Six elements out of ten. There's no scope signal, no conflict priority, and no first-line insight, because a bug in one function carries none of those risks.

Looking back, I would add two more sentences. The prompt has no ask-user gate, so if the cause had been somewhere other than the checkout function, the agent would have had to guess again. And the quality bar says nothing about behavior: "verify that one click submits exactly once" would have given the agent a concrete way to prove the fix works. It worked without them because the location and symptom were precise, but precision is what made the gap small, not the absence of the gap.

What changed

The difference was immediate:

First attemptSecond attempt
Prompt"The checkout page is not working. Detect and fix it."File, function, exact symptom, boundaries, quality bar
Files changedMultiple, most of them unnecessaryOne: src/checkout/page.tsx
The bugStill thereFixed
Side effectsYes, in unrelated parts of the appNone
What I got backA diff to untangleA focused diff plus a report explaining the fix

The prompt was longer, but the run was shorter. A few extra sentences up front replaced a wide search, a wrong fix, and a cleanup.

The bigger change was in how I work. I no longer treat a prompt as a wish. I treat it as a brief to a capable colleague who can't see my screen: what I saw, where it lives, what not to touch, and how we'll both know it's fixed.

Where the framework falls short

The framework helped, but it isn't complete, and it's worth knowing where it bends.

  • "Stop and ask" assumes someone is watching. For a long unattended run, nobody answers the question. A better instruction there is a safe default plus a visible flag: "leave it unchanged, add a TODO, and list it in the summary."
  • The role line does less than people think. "You're a frontend engineer" is the element people write first and the one that does the least work. The file paths and the symptom carry the weight. If you're dropping one element, drop this one.
  • Project instructions should carry the permanent rules. If your agent reads a CLAUDE.md or AGENTS.md, put your conventions and permanent boundaries there. A per-task prompt should hold only the facts of this task.
  • It covers what to add, not what to leave out. Pasting a whole log or a long chat history competes with the instruction for attention. Removing noise is part of context engineering too.

Summary

  • A vague prompt doesn't save time. It moves every decision to the agent, which fills each gap with a guess, and you pay for the guesses in tokens, side effects, and rework.
  • Most bad agent output is a briefing problem, not a model problem.
  • A good brief has up to ten elements, but almost no task needs all ten. Always include where and what, then add one element for each risk you can name.
  • For bug fixes, the elements that matter most are the exact symptom, the location, a boundary around code that must not change, and a way to verify the fix.

The shortest version fits on one line: say where, say what, say how to check it, then add one sentence for each risk you can name.