Releasing AI Changes: Evidence, Compatibility and Limits of Approval
Decide what evidence permits a model, prompt, retrieval or tool change to reach production. Compare release strategies, bind acceptance to tested revisions, preserve...
audience="Engineering leaders, application owners and release operators responsible for an AI-assisted business workflow." decision="Which changes need fresh evidence, which components must be tested together, and what acceptance permits during a controlled release." position="Accept a tested combination for a defined task and exposure scope. Component approval is useful evidence, but does not establish that the assembled workflow is safe to release." scope="A proposed release framework and a hypothetical service-request example. Not a customer result, certification or guarantee of model behavior." outputs={[ 'A change-impact and compatibility register', 'A versioned release evidence manifest', 'Separate behavioral and authority criteria', 'An acceptance validity policy', 'An in-flight work and rollback inventory', 'A release review worksheet', ]} />
Executive summary
An AI workflow can change without a new application binary. A source document changes, an index is rebuilt, a prompt is edited, a provider alias resolves differently, or an administrator enables another tool. A conventional deployment record may correctly identify the software version while leaving the behavior that matters to the user unexplained. When a bad recommendation or unauthorized write occurs, the team then cannot establish which combination was tested or whether the observed request belonged to that combination.
The proposed acceptance unit is a release manifest: the task contract, effective model configuration, prompts, retrieval controls, tool interfaces, executor policies and runtime revisions that together produced the reviewed evidence. This does not require every component to ship simultaneously. It requires evidence for compatible combinations and a way to identify which combination served each request. Acceptance remains limited to named users, tasks, data boundaries, actions and operating conditions.
The practical consequence is that a release approval and a transaction approval are different things. Permitting a new workflow revision does not authorize every action it proposes. A rollback can stop new use of that revision while leaving submitted operations unresolved. Operational and security evidence need separate verdicts: useful behavior, access boundaries, business execution and recovery cannot be reduced to one quality score. This paper explains how to choose that release model, what it costs to operate and when a smaller intervention is sufficient.
Scope, assumptions and the reader's decision
Assume an application helps an authenticated user understand a business request and may propose a change in a system of record. The application uses retrieved guidance and current operational data. It has an executor outside the model that can check identity, permitted actions and record revisions. Some changes are reversible; others can trigger notifications or external work that cannot simply be erased. These conditions make release compatibility and in-flight work consequential rather than administrative details.
The team must decide whether its existing deployment gate covers these changes. Ask whether a reviewer can reconstruct an affected request, whether source and permission changes are represented in evidence, and whether operators can stop new writes independently of stopping answers. If those capabilities already exist, extend their records rather than introduce another approval portal. If they do not, start with one consequential workflow and a limited acceptance scope. A company-wide governance policy without operational identifiers is unlikely to help an incident responder.
This framework assumes the business can name an accountable owner and define acceptable behavior. It does not supply industry-specific legal requirements, medical safety validation or independent certification. Nor does it assume a model can be made fully deterministic. It is a way to make a release decision explainable under uncertainty, with explicit restrictions. Organizations handling regulated decisions need qualified domain review in addition to the engineering evidence described here.
Distinguish behavior, compatibility and authority
Behavior is what the assembled application does for a task: which information it finds, how it interprets conflicting evidence, which answer or proposal it produces and how it handles uncertainty. Compatibility is whether its parts agree on their interfaces and meanings. An adapter may accept a new schema but interpret an omitted field differently. Authority is what the authenticated actor and executor are allowed to read or change. These properties overlap in an incident but require different tests and different owners.
A release record should not say only that the assistant passed evaluation. It should identify the assessed behavior, the compatible parts and the authority restrictions. A proposal can be accurate but unauthorized. A schema can be compatible while permitting an unintended business state transition. A model can decline one attack while another input reaches a privileged adapter. Keeping these statements separate lets the reviewer hold a release for one failed boundary without debating whether its overall answers became more fluent.
Treat application release acceptance as permission to expose a combination under stated conditions. Treat action authorization as permission for a particular operation against particular inputs. The release owner may accept a tool-capable workflow for a limited cohort, while the executor still requires the domain owner's approval before dispatching a consequential write. Neither record substitutes for the other. This separation also prevents a developer's successful sandbox run from becoming implied permission to operate on customer records.
Version the effective combination, not just the repository
Record the actual inputs that influence behavior. A useful manifest includes an immutable application artifact identifier, prompt revision, model identifier and available provider revision information, inference settings, retrieval configuration, source or index snapshot policy, output schema, adapter revision, executor policy and feature flags. Include the fixture set and grading rubric revisions that supported acceptance. These identifiers must be observable at request time, not reconstructed from what the team believes was deployed.
Some dependencies cannot be pinned exactly. A provider may expose only a model name, while a live source changes continuously. Record that limitation instead of inventing precision. Describe the observation window and the checks used to detect incompatible behavior. A release with a pinned prompt and an unversioned source is not reproducible in the same way as a replay against an archived source revision. Both can be operated, but they support different explanations and different confidence in a later comparison.
Do not store secrets in the manifest. Identify a credential policy or secret version through an approved reference, with restricted access to any sensitive details. Keep a readable release summary beside the machine-readable record. Operators need to know which tasks and actions are allowed without opening private prompts or source documents. The manifest should make the relevant evidence findable while preserving access controls on that evidence.
Map changes to their consequences
Classify a change by the paths it can affect, not by its file extension. A one-line prompt edit can alter tool selection. A source refresh can remove a restriction. A schema modification can change how a missing value is handled. A logging change may expose sensitive retrieved text even when the answer is unchanged. Small diffs are useful for diagnosis, but small size does not establish low consequence.
| Changed element | Consequence to investigate | Evidence that matters | |---|---|---| | Model or inference configuration | Different interpretation, abstention or tool choice | Repeated task fixtures, error classes and proposed actions | | Prompt or context assembly | Changed instruction precedence or omission handling | Before/after traces including supplied evidence | | Source, chunks or ranking | Different facts reaching generation | Source revision, retrieved spans and permissions | | Output schema or adapter | Changed defaults, units or business meaning | Structural checks and executor-level assertions | | Executor policy or identity mapping | Expanded read/write authority | Negative access cases and denied operations | | Runtime or feature flags | Different routing, retries or limits | Effective configuration and recovery observations |
The register is an impact hypothesis, not proof that every listed consequence occurred. The engineering owner marks which paths changed, the domain owner explains the business consequence, and the reviewer selects evidence accordingly. If a team cannot justify excluding a path, include it in the investigation or restrict exposure. This makes the argument visible instead of burying it in a generic low-risk label.
Diagram: evidence belongs to the assembled workflow
The diagram deliberately groups changes by consequence. Information controls determine which evidence is available. Generation controls influence the answer or proposal. Execution controls determine which operations may occur. A release can involve one or several groups. Their identifiers meet in the manifest because a reviewer needs to connect a result to the effective combination, not because the architecture requires a single service or a single deployment pipeline.
Keep the runtime architecture separate from this evidence view. A cloud inventory would not explain why a changed source and an unchanged model produced a different proposal. Conversely, this view does not establish network isolation, availability or data residency. Use it during release review to ask which identifiers are missing and which compatibility claim lacks a test. Those questions determine the next investigation more directly than adding another approval box.
Compare release options and their trade-offs
A fixed periodic review can be sufficient for a read-only, low-consequence internal assistant with stable sources and little change. It creates a predictable review workload and a clear opportunity to inspect accumulated changes. Its weakness is the gap between reviews. If a tool permission or source boundary changes mid-cycle, the schedule offers no immediate protection. Define exceptions that require review before exposure rather than treating the next meeting as the universal gate.
Per-component acceptance works when interfaces are stable and consequences are well-contained. A retrieval team can verify source eligibility and a model team can evaluate answer behavior. This preserves local ownership and avoids retesting unrelated work unnecessarily. Its weakness is assembly: individually accepted components can interact in an untested way. Require an explicit compatibility contract and representative integration evidence before treating component records as sufficient for a consequential workflow.
Manifest-based, consequence-led acceptance covers the assembled task. It can reuse component evidence while adding cross-boundary tests for the changed paths. The cost is maintaining dependency records, fixtures and reviewer decisions. That cost is justified when tool writes, sensitive access or material business decisions make interactions important. It is not justified by adding paperwork alone. Choose it because operators need to identify and control those interactions, then measure whether the records actually support diagnosis and safe release decisions.
Combined changes need interaction evidence
When a prompt and retrieval configuration change together, a better answer does not reveal which change caused it. A controlled comparison can run the old combination, prompt-only change, retrieval-only change and combined candidate against the same fixtures. This isolates an interaction that component evaluations might miss. The matrix is useful when the changes plausibly depend on one another, not as a requirement to test every theoretical permutation of the entire platform.
Suppose a prompt now asks for the latest policy while ranking now favors shorter passages. Either change alone may retain a qualifying exception; together they may omit it. The combined answer can look concise and helpful while authorizing a proposal the policy did not support. The reviewer needs the actual retrieved passages and proposed operation, not only a preferred-answer vote. If the result cannot be explained, holding the combined candidate is more defensible than accepting it because both component owners approved their parts.
Where a dependency cannot be exercised independently, record the limitation and treat the combination as the test unit. This may be necessary for a provider-managed feature or a tightly coupled schema/adapter migration. Explain which attribution questions remain unanswered and how exposure is restricted. A realistic integration result with explicit limits is more useful than a nominal component comparison that silently uses different inputs or cannot run the old contract.
Build evidence around task families
Define success for the task before evaluating a candidate. For an explanatory assistant, useful criteria may include finding the applicable source, acknowledging missing evidence and not revealing inaccessible information. For an action workflow, add recipient selection, revision checks, action limits and reconciliation. Keep mandatory prohibitions separate from preferences. A more readable explanation cannot offset a cross-tenant disclosure or an unauthorized write.
The Claude evaluation guidance recommends task-specific, measurable success criteria and evaluation methods suited to the criterion. This supports choosing different checks rather than treating one response similarity score as a release verdict. The proposal here adds release ownership and compatibility records; those are design recommendations, not claims that a provider's evaluation feature supplies an application governance system.
Organize fixtures by consequential families: ordinary supported requests, missing facts, conflicting sources, revoked access, stale approvals, malformed proposals and uncertain external outcomes. Report results for each family with its coverage and exclusions. Do not turn a handful of repeated happy-path runs into a claim about production accuracy. Record how examples were selected and which real cases are missing. Where a rare failure would be unacceptable, use a specific invariant and restrictive authority boundary rather than relying only on average observed performance.
Structural validity does not establish business validity
Schema validation is a useful first gate. It can reject a proposal with the wrong type, absent required fields or unexpected properties. The JSON Schema object reference documents these as separate constraints; merely listing properties does not require them or prohibit additional ones. Test the deployed schema and validator configuration rather than assuming a well-formatted example captures the contract.
After structural validation, evaluate meaning. Does the selected record belong to the actor? Is the amount expressed in the expected currency? Does an empty list mean leave unchanged or clear existing values? Is the proposed transition permitted from the current state? These checks belong in application and domain logic. A schema may represent them partly, but syntax alone cannot establish current record ownership or authority to perform the operation.
Include the adapter in release evidence whenever it maps model output into an API call. Otherwise a candidate can pass a model-only evaluation while the adapter supplies a dangerous default. Preserve both the proposal and the normalized operation in test traces. A reviewer should be able to see whether the system rejected an invalid value, requested clarification or quietly substituted something else. Silent repair is itself behavior that needs acceptance when it changes a consequential field.
Access changes are behavior changes
Retrieval tests need identity-specific cases. If a user loses access to a source, an old cached answer or conversation may still contain its contents. A release test that queries only as an administrator will not detect the problem. Decide where authorization is enforced, which metadata it uses and how changes reach indexes, caches and replay artifacts. The release record should describe these boundaries without claiming that a single filter guarantees confidentiality across the application.
The Azure AI Search access-control overview describes document trimming based on identity and indexed permission metadata, with different approaches and availability conditions. The general lesson for this framework is that data eligibility is part of retrieval behavior. Do not infer that native controls are available for every source, API version or identity model. Verify the chosen implementation separately.
Acceptance can permit a stable application to read changing live sources if the source admission and permission processes have their own controls. It should not imply every new document has been reviewed as an instruction. Test the source boundary with disallowed content and revoked identities. Retain enough metadata to explain what was eligible at the time, subject to retention policy. A forensic archive that ignores deletion obligations creates a different risk and is not a free solution to reproducibility.
Security evidence must reach the executor
An indirect instruction in retrieved material can affect an answer or proposed action. The OWASP prompt-injection guidance identifies this class of risk and recommends layered controls, restricted privilege and adversarial testing. It does not establish a foolproof prevention method. A release review should therefore ask what the executor refuses even if the model produces a persuasive or correctly shaped request.
Use fixtures that place malicious instructions in the sources the application actually reads, including tool results and attachments where relevant. Inspect the resulting proposal and the executor's response. Did the model request a prohibited operation? Was it denied? Was sensitive information already included in the answer before denial? These are distinct outcomes. A blocked write is meaningful evidence for the write boundary, but does not by itself prove the answer boundary remained safe.
Keep authorization outside a free-form interpretation of model text. A natural-language statement that the user approved something is not an approval record. The executor needs authenticated scope, significant operation data and an enforceable state transition. Review changes to privilege policies as consequential releases even when the model and prompt are unchanged. Otherwise the same previously harmless proposal can become executable without a new behavior test.
Acceptance has a validity boundary
Record what would invalidate release acceptance. Changes to a mandatory policy, permitted action, output meaning or required source boundary usually require new relevant evidence. Minor source additions may be covered by an accepted ingestion process if its constraints remain unchanged. Define that distinction before the change arrives. An approval that survives every possible configuration edit is not meaningfully tied to tested behavior.
Validity is not only a date. A release may remain within its accepted conditions for a long period, while another becomes invalid immediately after an access or adapter change. A review interval can prompt reassessment, but it cannot detect all invalidating events. Combine event-triggered checks with time-based review appropriate to the workflow. State the restricted mode used when required evidence or ownership is unavailable, including whether read-only answers remain permitted.
For individual transactions, OWASP's transaction-authorization guidance emphasizes server-side enforcement, significant transaction data and limited validity. Applying that distinction here means a release-approved system must still recheck a particular operation's authority before execution. The paper does not prescribe a universal expiry duration or authentication method. Those choices depend on the action, exposure, identity model and domain requirements.
Diagram: stopping exposure is not undoing work
The view separates work by its state when exposure stops. New requests can be routed away from a revision. Queued proposals can be held for a fresh compatibility and authority check. Dispatched operations may still commit after the application stops waiting. Committed effects require a domain-specific recovery decision. Collapsing these into one rollback button hides the cases that matter most during an incident.
Identify operations by stable business references and preserve their originating manifest. A request can outlive the release that created it. If an operator replays it through a newer adapter, the meaning may change. Either continue it under a still-valid compatible contract or explicitly convert and reauthorize it. Do not treat changing the deployment version as permission to reinterpret an unresolved business instruction.
Design rollback before permitting writes
Maintain a last accepted combination, not merely a previous container image. A newer index schema, adapter contract or persisted proposal may be incompatible with the old binary. Rehearse which components can be restored, which must remain forward-compatible and which require a migration. If the rollback depends on an unavailable source snapshot, identify that limitation before accepting the release rather than during the incident.
For externally submitted work, restoring software is not recovery. An email may already have been sent, a payment request accepted or an appointment changed. Stop further dispatch, identify affected operations and reconcile authoritative outcomes. Compensation may be possible, but it is another business action with its own permission and consequence. Do not automatically reverse every affected record or retry every uncertain operation after switching versions.
The companion AI action recovery paper examines this uncertainty in detail. In the release framework, its main implication is an acceptance prerequisite: an operator must know how to find unresolved actions and who can decide their disposition. A release should be held if it enables writes whose outcomes cannot be identified after a timeout. Better answer evaluations do not repair that operational gap.
Use shadow and canary modes for different questions
A shadow run can compare outputs without serving the candidate's result, provided its tools and downstream effects are genuinely isolated. It is useful for discovering changed retrieval, abstention or proposed actions. It does not prove that users will behave the same way when they see the candidate, or that real writes will recover correctly. Avoid sharing production write credentials with a supposedly observational run. Inspect connectors and callbacks, not just the feature flag name.
The Google SRE canary chapter describes limited exposure and comparison with a control to inform rollout. Applying that approach to an AI workflow requires suitable observations: unsupported proposals, denied actions, unresolved operations and access failures may matter more than response latency alone. The thresholds and exposure plan are application-specific proposals here, not numerical recommendations taken from Google's operational experience.
A canary for a tool-writing system is real exposure. Define the cohort, action limits, stop authority and recovery route before activation. Do not let a candidate write the same business operation as a control in an attempt to compare them. Use synthetic isolated writes for paired testing and authoritative outcome observation for permitted production operations. Where consequences cannot be contained, keep the experiment read-only until a safer comparison is available.
Worked example: changing a service-request assistant
Consider a hypothetical internal assistant that reads service policies, summarizes a request and proposes a scheduling change. The executor can create a draft, but confirming a customer appointment requires separate permission. The team wants a clearer prompt and a retrieval change that reduces irrelevant passages. This example is a design exercise, not an Ampity customer deployment or a measured improvement claim.
The change register identifies two coupled paths: policy selection and proposal construction. The fixtures include a permitted request, an unavailable slot, conflicting policy revisions, an inaccessible customer record and an already-confirmed appointment. The team compares both single changes and the combined candidate. It preserves retrieved revisions, proposals, adapter output and denied operations. A more concise answer is treated as a preference; selecting an inaccessible record or bypassing confirmation authority is a release-stopping failure.
Assume the combined candidate handles ordinary requests well but omits a qualification when two sources conflict. The correct decision is not to average that failure into the ordinary result. Hold the affected task family, revise the source-selection or clarification behavior and rerun its evidence. If the owner chooses to expose only unrelated read-only tasks, record that narrower acceptance explicitly. The candidate is not accepted for scheduling changes simply because part of its workload remains useful.
Account for review capacity and operating cost
The release process consumes engineering time, domain review, model calls, test storage and incident attention. Estimate those costs by changed task family rather than by counting releases alone. Reusing a stable negative-access fixture may be cheap; adjudicating ambiguous business proposals can be expensive. Automated checks should handle exact invariants while qualified reviewers address meaning. Automating a subjective grader without testing its reliability can move effort into hidden false acceptance.
As an illustrative planning calculation, suppose a candidate produces sixty cases requiring two minutes of manual adjudication each. That is two hours of direct review before disagreements, reruns and record preparation. If the team has only one hour available, the answer is not to waive half the cases silently. Reduce scope, schedule sufficient review or improve the fixture rubric so more cases can be checked deterministically. These figures are assumptions for capacity planning, not a benchmark or a suggested universal sample size.
Measure queue age for release reviews, repeat investigations caused by missing traces and incidents attributable to unrecorded configuration changes. A manifest process that creates delay without improving those outcomes needs revision. Conversely, rapid approvals obtained by merging security failures into a favorable average are not evidence of efficiency. The economic objective is to avoid wasted or harmful exposure while keeping justified changes executable with proportionate evidence.
Assign ownership without making every reviewer responsible for everything
The engineering owner describes changed dependencies and supplies reproducible traces. The domain owner defines permitted business outcomes and ambiguous-case judgments. The security owner reviews access and privilege boundaries where changed. The release operator verifies the effective manifest, exposure limits and stop mechanism. One accountable acceptance owner reconciles their findings and records the final scope. This is a responsibility split, not a requirement for five separate people in a small team.
State who can stop exposure without waiting for a committee. Give that person or automation a tested path to disable dispatch and retain unresolved-operation records. Separately identify who may approve recovery actions. An incident operator's authority to stop a release should not automatically include permission to alter financial or customer records. Where the same person holds several roles, preserve the distinction in the record even if the meeting is short.
Avoid sign-off by silence. An absent domain reviewer leaves a gap; it does not imply acceptance. If the approved scope excludes consequential writes until that gap closes, make the restriction enforceable in the executor and visible to users. Ownership should reduce ambiguity for the operator, not merely create names on a document. A release decision is complete only when its conditions can be enforced and its evidence can be found by the people responding to it.
Limitations and when a lighter process is enough
A manifest cannot make an unknown provider revision reproducible, prove unseen adversarial cases safe or determine legal suitability. It can record those gaps and restrict exposure accordingly. Evaluation sets also age: task distributions, sources and user expectations change. Review coverage after meaningful incidents and observed shifts, but do not automatically turn private production conversations into fixtures. Approve retention, redaction and access before reusing their contents.
A low-consequence read-only assistant with a stable source set may need only an extended deployment record, representative answer checks and an access review. It may not need a separate release service or elaborate compatibility graph. The same team should reconsider that choice when enabling writes, adding sensitive sources or changing authority. Complexity follows consequence and diagnosability, not the presence of the letters AI in a product name.
For safety-critical or regulated decisions, this framework is incomplete without the applicable assurance process and qualified domain expertise. Restrict model participation where evidence cannot support the consequence. In other settings, the smallest effective improvement may be exposing source revisions, rejecting ambiguous proposals or introducing a dispatch stop. Prefer a control the operator can demonstrate over a broad policy that the runtime cannot enforce.
Release review checklist and next steps
Use the following worksheet on one upcoming change. It is a decision artifact, not a substitute for its linked evidence. An unknown answer means investigate or reduce accepted scope. Record the reviewer and the evidence location beside each conclusion so another operator can reconstruct the decision later without relying on the original meeting participants.
| Review question | Required evidence | Hold condition | |---|---|---| | What task and users are accepted? | Scoped task contract and exposure plan | Scope is implicit or broader than fixtures | | Which combination was tested? | Effective manifest and request traces | Unidentified dependency changes | | Which behavior changed? | Per-family comparison and exclusions | Critical failures hidden in an average | | What stays forbidden? | Access and executor negative cases | Model text can grant authority | | When does acceptance expire? | Invalidating events and restricted mode | Policy edits silently inherit approval | | What happens to pending work? | Queue inventory and revision checks | Old proposals are replayed unreviewed | | What cannot rollback undo? | Dispatch ledger and recovery owner | External outcomes cannot be reconciled | | Who can stop exposure? | Rehearsed stop path and accountable owner | Stopping depends on an unavailable reviewer |
Start with the AI change-regression release-gate playbook to build the execution evidence. The narrower articles on prompt-change tests and retrieval changes explain the diagnostic comparisons. Do not expand the governance scope until this first workflow has an identifiable manifest, an enforceable acceptance boundary and a demonstrated treatment of unresolved work.
If you want Ampity to help review these boundaries, the relevant work is a scoped production AI systems assessment or agentic workflow implementation. The useful engagement output would be a dependency register, test evidence, permitted-action contract and recovery procedure agreed for your application. Reading and downloading this resource does not require an enquiry. Any contact request is your choice.
Primary references
The following pages were inspected for this local draft on 4 October 2026. They support the specific evaluation, schema, access, security and canary concepts cited above. The release manifest, ownership model, diagrams and hypothetical examples are the paper's proposed framework, not vendor-endorsed requirements.