What Should a Prompt-Change Regression Test Catch?
Test a prompt revision against task meaning, evidence use, uncertainty, tool arguments and failure paths. Build a regression record a release owner can inspect.
Test the behavior the application depends on
A prompt-change regression test should catch changes to the meaning of the task, evidence use, uncertainty handling, output interpretation and proposed tool actions. It should also check failure paths the application relies on. A response can remain fluent and valid JSON while proposing the wrong action or concealing missing evidence.
Start with one application contract. For a document assistant, that might be preparing a draft from supported fields while holding unresolved values for review. For a service assistant, it might be answering from approved material and collecting an enquiry only when the visitor asks for contact. Those contracts require different fixtures and different release decisions.
The method below is a proposed engineering practice. Its examples are synthetic, not Ampity customer results or evidence that a particular prompt is secure. The tests support a release decision within a defined scope; they cannot prove correct behavior for every future input. Keep ordinary authorization and validation controls outside the prompt.
Write the expected behavior before editing the prompt
Suppose a team changes an instruction from “prepare an invoice draft with unresolved fields clearly marked” to “complete the invoice efficiently.” On a clean sample, both prompts return a plausible draft. On a sample without currency evidence, the new prompt fills in the currency from context. The response parses correctly, yet the workflow has lost its hold condition.
Write the expected behavior for that fixture before reviewing candidate outputs. The application may accept an unresolved currency in a draft, provided the uncertainty is visible and posting remains blocked. Alternatively, it may refuse draft creation until currency is supplied. The test needs the actual policy; it should not let the grader choose whichever answer sounds helpful.
Separate what must remain unchanged from what the revision intends to improve. A clearer explanation of a held field is an intended change. Allowing that field to reach a posting proposal is a regression. Record the distinction so a reviewer can approve a wording improvement without unknowingly widening operational scope.
Anthropic's evaluation guidance recommends task-specific criteria and measurable evaluation, including edge cases. The fixture policy proposed here applies that principle to the behavior an application consumes.
Freeze the rest of the experiment
Record the baseline and candidate prompt revisions, model identifier, request settings, tool definitions, retrieval snapshot, application adapter and grading rules. Preserve the effective prompt after template substitution. A source file alone may omit the instructions, defaults or contextual values added by the application at runtime.
Run both candidates against the same approved fixture inputs where isolation is possible. If the model or corpus changed as well, label the run as a combined change. It can still be useful release evidence, but it does not attribute a failure specifically to the prompt. Do not describe a changing provider response as a controlled experiment when other inputs also drifted.
Keep secrets and unnecessary personal data out of fixtures. Synthetic examples are useful for known boundaries, but they do not establish representative performance on their own. Where approved historical examples are available, minimize them and preserve the context needed to judge the task. Record which input populations the test set does not cover.
Pin the grader configuration too. Otherwise the same saved candidate response can receive a different verdict after a grader update, and the apparent regression may come from evaluation rather than application behavior. Store the rubric and the reason for each consequential verdict alongside the result.
Inspect the contract at several layers
Use deterministic assertions where the contract is exact. A response with an unrecognized status, a missing required key or an invalid identifier can fail a parser check. Then review whether the valid values mean what the workflow expects. A status of “ready” needs evidence of readiness, not merely membership in an allowed enum.
| Test layer | Synthetic fixture | Expected evidence | Example regression | | --- | --- | --- | --- | | Task interpretation | Visitor asks a technical question without requesting contact | Answer supplied without a mandatory lead form | New prompt requires email before answering | | Evidence use | Invoice lacks supported currency | Unresolved field and applicable hold remain visible | Candidate infers a currency and labels the draft ready | | Tool proposal | User asks to inspect an existing record | Proposed arguments remain read-only | Candidate proposes a record update | | Output contract | Supported draft can be prepared | Defined schema and correct field semantics | Valid JSON places a source reference in an account field | | Untrusted content | Retrieved text contains conflicting instructions | Content remains evidence, not execution authority | Candidate follows embedded instructions | | Dependency failure | Tool returns an unavailable result | Honest delay or defined alternative path | Candidate states the task completed |
Inspect the application's interpretation of the response. A test that stops at model text can miss an adapter that treats every nonempty output as accepted. Run the candidate through the same parsing and routing logic used by the application, with effectful destinations replaced by isolated test doubles.
Keep the test doubles faithful to the behaviors under examination. A connector stub that always returns success cannot test unavailable or ambiguous outcomes. Include a recorded response shape for each relevant failure, and make the expected application state part of the assertion.
Include adversarial and awkward inputs
Test untrusted instructions inside documents and retrieved snippets, including content that claims to override the application's rules. OWASP's prompt-injection guidance describes both direct and indirect injection and recommends adversarial testing and privilege controls. Passing a few such fixtures does not establish immunity.
Add ordinary awkward cases alongside adversarial ones: incomplete requests, ambiguous references, contradictory evidence and a conversation that changes topic. These cases test whether a shortened prompt has removed behavior needed to ask a useful clarification or stop at the right boundary.
A refusal is not automatically the correct result. If the assistant can safely explain available information without performing an unsupported action, an unnecessary refusal may reduce usefulness. Grade the permitted response for the specific fixture rather than awarding points for the presence of cautious language.
Exercise conversation history. A previous user's preference may help resolve a presentation choice, but it should not become authorization for a later record change. Inspect whether the candidate prompt carries forward the right context and discards obsolete scope. A single-turn test set leaves that behavior unexamined.
Judge repeated runs without hiding variability
Where responses vary across repeated runs, preserve that variation. Set a repetition budget appropriate to the decision and report its size. One successful run demonstrates one observed result, not consistent behavior. Conversely, identical formatting across runs does not establish stable task meaning.
Use exact checks for invariants, such as the absence of an unauthorized tool proposal. For nuanced answer quality, use a written rubric and review disagreements. If a model grades responses, compare a sample of its judgments with qualified human review before relying on that grader to screen consequential changes.
In a hypothetical test set, a candidate may improve 18 explanatory responses while introducing two unsupported completion claims. An average preference score can make that revision look attractive. The release owner needs the two claims and their consequences before deciding whether improvement justifies release. Hard stop conditions should stay visible outside aggregate scores.
Record which failures existed in the baseline and which were introduced. A baseline failure still needs ownership, but it answers a different question from a newly introduced regression. Preserve both in the report; removing known failures from the fixture set can make future comparisons look cleaner without improving the application.
Make every failed fixture reproducible
A compact regression record contains the case identifier, input revision, expected behavior, baseline trace, candidate trace, grading rule, observed difference, impact and owner. Include enough source evidence to review the verdict without copying an entire sensitive conversation into a general-purpose report.
For the missing-currency example, retain the source revision, unresolved-field state and proposed next action. The useful failure description is “candidate proposed posting with an inferred currency,” rather than “answer quality decreased.” The former identifies what engineering must change and what a rerun must prove.
When the intended behavior changes legitimately, create a reviewed fixture revision with the reason and policy owner. Do not overwrite the expected output solely to make the candidate pass. The resulting evidence should distinguish an accepted contract change from a correction to a test that was wrong.
Maintain held-out cases that are not used for prompt tuning. Frequent iteration against the same examples can teach a team to improve those examples while leaving the wider task unchanged. Use fresh, independently reviewed cases to test whether the proposed revision generalizes within the agreed input population.
Leave a release decision the next engineer can inspect
The release record should identify the exact candidate, accepted scope, stop conditions, unresolved failures and rollback reference. A prompt rollback needs the prior effective instructions and compatible application configuration. Reverting a sentence while keeping a changed tool adapter may not restore the earlier behavior.
Before release, choose one failure fixture and walk it through the application with the owning team. Confirm that it produces the expected hold, clarification or denial without a real customer effect. Retain the trace as acceptance evidence, and name who can disable the revision if live observations contradict the test result.
The AI action recovery playbook covers effectful tool failures beyond prompt testing. Read why retrieval changes break AI answers when the evidence supplied to the model has changed. If you want Ampity to examine a release boundary, share the task and its constraints. The technical guidance remains available without providing contact information.