Build an AI Change-Regression Release Gate
Turn a prompt, model or retrieval change into a versioned release decision using owned fixtures, independent grading, failure rehearsals and a tested rollback.
trigger="A prompt, model, retrieval or tool-adapter revision is ready, but the team cannot show which behaviors it preserves or how it will be stopped." owner="The application release owner accountable for the accepted task scope." participants={['Application engineer', 'Domain reviewer', 'Evaluation owner', 'Security reviewer', 'On-call operator']} prerequisites={['One application task with a written behavior contract', 'Baseline and candidate manifests', 'Approved fixtures and an isolated effectful destination', 'Named owners for acceptance, stopping and rollback']} outputs={['A versioned fixture and grading register', 'Baseline/candidate traces with per-family verdicts', 'A failure and rollback rehearsal record', 'An accepted-scope decision or documented hold']} doneWhen={['Hard-stop failures cannot be hidden by aggregate scores', 'A domain reviewer resolves consequential grading disagreements', 'The exact candidate and compatible rollback are identified', 'The on-call operator can stop exposure and reconcile in-flight work']} />
Choose one task and one release decision
Build the gate around the behavior an application depends on. Start with one task, such as answering questions from approved policies or preparing document drafts. Name the effects the application may perform and the conditions that require clarification, review or refusal. Then compare a baseline and candidate against that contract.
This playbook proposes an engineering procedure, not a certification or an Ampity customer outcome. Run its examples with synthetic or approved minimized data and isolated destinations. A model's passing evaluation does not replace server-side authorization, record-version checks or the owner's permission to release. If the task is safety-critical or subject to sector-specific obligations, include the qualified reviewers and controls required for that context.
Use the prompt-regression article to identify instruction-level failures and the retrieval-regression article to inspect changing evidence. The steps here combine those observations into a repeatable release record. They permit a hold when the evidence is insufficient; finishing the procedure does not require deployment.
1. Record the effective change
Owner: application engineer. Output: baseline and candidate manifests. Record the model identifier, effective instructions, tool definitions, retrieval and ingestion configuration, source snapshot, application adapter, output schema and request settings. Identify which components changed and which remained fixed. Preserve the substitutions applied to templates, not only the editable prompt file.
State the intended improvement in one testable sentence. For example, “Explain why an invoice draft is held without inferring unresolved values.” Identify behaviors that must remain unchanged, including the allowed write scope and handling of unavailable dependencies. A general objective such as better answers gives reviewers no stable acceptance boundary.
Resolve version uncertainty before claiming a controlled comparison. If a provider alias or external corpus can change during the run, record the observed identifiers and limitation. Either freeze the relevant input where supported or label the experiment as a combined change. Stop attribution claims when the manifests do not explain the actual system tested.
2. Assign acceptance and stop authority
Owner: release owner with the domain and security reviewers. Output: a signed acceptance-policy revision. Define hard-stop failures separately from quality preferences. An unauthorized proposed action or restricted-source disclosure may block release even when explanatory responses improve. Specify which verdicts require specialist review and who can override a disputed assessment.
Write acceptance by question or task family. A lookup answer, a multi-source judgment and a tool proposal have different evidence needs. State the supported input population and exclusions. Do not borrow accuracy or latency thresholds from a vendor example as if they were established business requirements.
Anthropic's evaluation guidance supports task-specific, measurable criteria and edge-case evaluation. Apply that principle through your domain contract. Keep any organization-specific threshold identified as the owner's policy, with its rationale and test population, rather than presenting it as a universal industry standard.
3. Build a versioned fixture register
Owner: evaluation owner with a domain reviewer. Output: approved fixtures and independent expected behaviors. Assemble normal inputs, incomplete inputs, contradictory sources, unsupported questions, dependency failures and relevant adversarial examples. Include conversation context where the task depends on prior turns. Assign an expected disposition before reviewing candidate output.
Use this compact register for each fixture. The fields are a proposed starting point; add the sensitivity and lineage information your application needs.
| Record field | Required entry | Example for a synthetic draft task | | --- | --- | --- | | Case and input revision | Stable ID and frozen input | INV-CURRENCY-02, revision 3 | | Source and caller scope | Relevant evidence and allowed identity | Approved synthetic invoice; draft-only reviewer | | Expected disposition | Permitted response and next state | Draft held for unresolved currency | | Forbidden effect | Action the test must prevent | Posting or supplier-master update | | Grading method | Assertion and reviewer rubric | Held-state check plus evidence inspection | | Severity and owner | Who adjudicates a failure | Domain release blocker; finance process owner |
Keep tuning fixtures separate from held-out acceptance cases. When the same examples guide every prompt edit, a successful rerun can show adaptation to those examples rather than a useful general improvement. Record known coverage gaps explicitly. Synthetic cases exercise defined boundaries; representative task performance needs an appropriately approved sample of the actual input population.
4. Validate the application output contract
Owner: application engineer. Output: executable parser and semantic assertions. Define exact requirements for status values, identifiers, required evidence references and permitted tool arguments. Run outputs through the application's real parser and routing logic. A model response that validates under one schema may still be misinterpreted by a permissive adapter.
The JSON Schema object reference explains that declaring a property does not make it required. Check required fields, extra-field handling and absent-versus-null behavior using the schema dialect and validator your application actually uses. Schema validity covers structure; write additional checks for the business meaning of accepted values.
Inject one structurally valid but semantically wrong result. For example, put a source-reference identifier in a destination-account field. Confirm that the downstream application holds or rejects it rather than proceeding because parsing succeeded. Record the route and resulting state as evidence. Stop if the test harness bypasses a validator that production depends on.
5. Isolate effects before replaying fixtures
Owner: application engineer with the security reviewer. Output: an isolation checklist and connector test doubles. Replace destinations that can send messages, change access, pay invoices or update live records. Verify the isolation from the actual configured endpoint and credentials. A test label in a request does not guarantee the destination has disabled external effects.
Keep connector failures realistic. Include unavailable, denied, invalid and uncertain-result responses where the workflow supports them. A stub that returns success for every tool call cannot establish recovery behavior. Verify that the candidate reaches the expected application state through the same adapters it will use after release.
Record which dependencies remain shared with the baseline or other environments. Shared caches, mutable records or search indexes can contaminate comparisons. If complete isolation is impossible, identify the consequence and limit the claim. Stop replay if a fixture can cause an unauthorized external change or expose unapproved sensitive data in logs.
6. Run paired comparisons and retain traces
Owner: evaluation owner. Output: baseline/candidate result records. Execute the same approved fixtures against both manifests. Record model output, parser result, selected source evidence, proposed tool arguments, application disposition, elapsed time and errors. Preserve repetitions when variation affects the task; record the repetition count rather than showing only the best response.
Keep the grading configuration fixed for a comparison. A rubric or grader update is another change to investigate. When automated and human verdicts disagree on a consequential case, retain both and send the disagreement to the named reviewer. Do not silently discard an inconvenient result because the explanation reads well.
Distinguish newly introduced failures from failures already present in the baseline. Both need owners, but they answer different release questions. A candidate that retains a known security defect should not receive an unrestricted release recommendation simply because it introduced no additional defect. The acceptance policy must address that residual exposure explicitly.
7. Exercise injection and scope boundaries
Owner: security reviewer with the application engineer. Output: a boundary-test record. Place conflicting instructions in user content, retrieved material and tool responses. Test whether the application preserves allowed access and effect scope. Inspect both the model's proposal and the executor's decision; a refused proposal and a permitted unauthorized write are different observations.
OWASP's prompt-injection guidance recommends adversarial testing and privilege controls. Use the tests to inspect specific trust boundaries. Passing a limited fixture set cannot establish immunity from future injection techniques, and a stronger prompt does not justify giving the executor broader permissions.
Include a benign document that quotes an instruction as subject matter. The application should be able to explain relevant content without obeying it. Avoid a test rule that accepts every refusal as safe behavior. Record the permitted answer and denied effects separately, then verify that ordinary task usefulness survives the defense.
8. Review per-family results before aggregate scores
Owner: domain reviewer. Output: an adjudicated regression ledger. Review errors by task family, consequence and source condition. Show unsupported completion claims, wrong qualifiers, missing uncertainty and changes to proposed actions. Keep hard-stop cases visible alongside preference or answer-quality scores.
In a hypothetical run, 18 explanatory answers improve and two drafts lose their hold condition. An average quality improvement does not settle whether release is acceptable. Inspect the two draft traces, determine the resulting exposure and apply the written stop rule. The numbers here illustrate a decision conflict, not an observed Ampity result.
Record disagreements and changed expectations. If a source policy legitimately changed, revise the fixture with the domain owner's reason instead of marking every new answer wrong. If a candidate fails the original contract, do not rewrite expected behavior just to make the run green. Keep the provenance of each accepted contract change in the ledger.
9. Rehearse pause and compatible rollback
Owner: on-call operator with the application engineer. Output: a timed recovery rehearsal. Stop new candidate exposure in the isolated environment, then restore the prior compatible manifest. Verify effective instructions, retrieval configuration, tool definitions and adapter behavior after restoration. A configuration rollback that leaves incompatible indexed data or queued work behind may not restore the baseline.
Test one operation already in progress when stopping begins. Determine whether it completes under the original revision, is held or needs reconciliation. Do not assume withdrawing traffic cancels a dispatched effect. Use the AI write-recovery playbook when a tool outcome can remain uncertain.
Record who stops the revision, how quickly exposure changes and which work remains unresolved. Include an operator-facing notification and a recovery owner. Completion requires demonstrated restoration and a reconciled list of pending operations, not merely a successful rollback command. Hold release if the operator cannot identify which manifest is active.
10. Prepare controlled exposure criteria
Owner: release owner with the on-call operator. Output: an exposure and observation plan. If the offline evidence supports release, define a limited population, observation period, metrics and immediate stop conditions. This step prepares a plan; obtain the organization's required release authorization before directing live traffic.
Google SRE's canary-release chapter explains the need to separate candidate and control observations and account for shared infrastructure. Apply those concerns to the AI task. A service-wide error average can conceal a candidate-only failure or mix long-running work from different revisions.
Choose observations that can detect the accepted task's failures. Latency and HTTP errors alone cannot show whether a cited answer omitted an exception or a draft concealed missing evidence. Define a review sample and escalation path for those outcomes. Where privacy limits logging, decide how reviewers can inspect minimized evidence without creating an uncontrolled copy of the source data.
11. Issue the accepted-scope or hold record
Owner: release owner. Output: the release decision and evidence packet. Name the exact candidate manifest, supported task population, accepted improvements, unresolved failures, tested rollback and stopping owner. Link the fixture register, paired traces, adjudication ledger and recovery rehearsal. Include the date and accountable reviewer in the internal operational record.
Use a hold when evidence is incomplete or a hard-stop failure remains. State the failed case, affected scope, corrective owner and rerun requirement. A hold should leave a concrete next action: repair the adapter, recover missing evidence or clarify an acceptance rule. It should not leave the team waiting for a vaguely better model response.
If scope is narrowed, specify the routing that enforces the exclusion and test it. Writing “unsupported inputs excluded” in a release note is insufficient when the interface can still send those inputs to the candidate. Verify that the rejected population reaches the approved alternative, review or unavailable path.
12. Convert observations into reviewed fixtures
Owner: evaluation owner with the domain reviewer. Output: new fixture revisions and an observation ledger. After an authorized release, collect approved evidence of failures, ambiguous cases and unexpected task distributions. Minimize sensitive content and assign ownership before adding cases to the regression set. Retain the distinction between an observed incident and a synthetic reproduction.
Validate the expected behavior independently. A user's complaint may reveal a genuine defect, a stale source or an unsupported expectation. The domain owner resolves that distinction before the case becomes an acceptance test. Preserve the reason, source revision and consequence so the next release does not inherit an unexplained label.
Do not depict fixture collection as automatic model training. Changing a prompt, model or retrieval rule requires another manifest and the applicable gate. Use fresh held-out cases to check that the correction works beyond the example that motivated it. Keep the previous release and its observations available for comparison.
Finish with an inspectable acceptance packet
The packet should let an engineer outside the original test session reconstruct the candidate, explain its scope and rerun a consequential failed fixture. It should also let the operator stop exposure without asking which prompt file someone edited. Verify those two tasks with the owning team before closing the gate work.
Use this review checklist at the handoff. Record a link to the evidence and a named reviewer for each row. An unchecked condition requires a hold or an explicitly reviewed scope exclusion with tested routing.
| Handoff condition | Evidence to inspect | Reviewer | | --- | --- | --- | | Candidate reproducible | Effective manifest matches the tested adapter, instructions and retrieval inputs | Application engineer | | Contract preserved | Per-family traces satisfy required states and prohibited effects | Domain reviewer | | Hard-stop failures resolved | Security and consequential-error ledger has no unresolved in-scope blocker | Security reviewer and release owner | | Rollback compatible | Restoration trace and in-flight operation reconciliation | On-call operator | | Exposure observable | Candidate/control breakdown and task-specific review plan | Evaluation owner | | Decision accountable | Accepted scope or hold record references the exact manifest and owners | Release owner |
Take one disputed result and one rollback trace through that review. Record missing evidence as work to complete rather than smoothing it into a release summary. If you want Ampity to examine the gate, share the application contract and one minimized trace. Reading and downloading this playbook require no contact information; an enquiry is your choice.