Review AI-Generated Code Before It Reaches Production

Review one AI-assisted change against its behavior contract, access boundaries, dependency tree and recovery evidence before accepting the exact release artifact.

trigger="An AI assistant has proposed a working-looking change, but the release owner does not yet have evidence for the behavior and risks introduced by its actual diff." owner="The engineering lead accountable for the accepted change scope and release evidence." participants={['Change author', 'Independent reviewer', 'Security reviewer', 'Dependency maintainer', 'Release operator']} prerequisites={['An agreed behavior contract and exclusions', 'A pinned base and candidate revision', 'An isolated environment without production authority', 'Independent expected results and a recovery procedure']} outputs={['A diff and risk register', 'Verified dependency and configuration changes', 'A version-bound test record', 'A recovery rehearsal', 'An accepted-scope or hold decision']} doneWhen={['Every changed boundary has proportionate evidence', 'Required failures and denied access are tested', 'Dependencies and build identity are verified', 'Open risks have explicit dispositions', 'The release artifact matches the reviewed candidate']} />

Accept the change, not the assistant's confidence

Review AI-generated code by inspecting the complete proposed change, defining expected behavior independently and collecting evidence for the boundaries it alters. Treat explanations, suggested tests and automated review comments as inputs to the review. A named engineering owner accepts an identified candidate and its limitations; the assistant does not grant release authority.

Consider a hypothetical document service. An assistant adds asynchronous CSV export, a convenience library and a download endpoint. A local demonstration produces a file, but the change also creates a background job, selects tenant data and stores an artifact that remains available after the request ends. Each of those paths needs review. A successful demonstration of one authorized user's export does not establish that another tenant cannot retrieve it or that a cancelled job stops publishing files.

This procedure uses synthetic tenants and inert storage and notification adapters. It does not report a customer incident, measure the quality of a particular model or certify generated code as safe. The depth of review follows the changed behavior and consequence. A small copy edit needs different evidence from an export worker with access to personal records. Production writes, credential changes and deployment remain separately authorized actions.

1. Write the behavior contract before reviewing implementation

Owner: engineering lead with product owner. Output: accepted behavior contract. State the user operation, permitted actor, data scope, expected result and failure response. For the export fixture, define which records may be included, when permissions are checked, how a user discovers completion and how an unavailable result is presented. Decide whether access withdrawal must prevent a queued job from running, prevent retrieval or both, using the product's actual policy.

Record exclusions and nonfunctional limits. An export request should not quietly include archived records, change data retention or email a file unless those behaviors are approved. Choose workload limits from the product's requirements and operational budget rather than an invented universal file-size threshold. Include the maximum simultaneous work, handling of oversized requests and the consequence of rejecting an operation after partial processing.

Give each requirement an observable acceptance result and an owner. The expected CSV contents should come from the synthetic input and the agreed contract, not from reading the candidate's output and accepting it as the reference. If the team cannot state what the operation should do under denied access or interruption, resolve that requirement before treating a clean test run as sufficient.

2. Pin the base, candidate and scope of the diff

Owner: change author with independent reviewer. Output: revision-bound change inventory. Identify the actual base commit, candidate commit, dependency lockfile, build configuration and relevant generated artifacts. Inspect added, deleted and renamed files, workflow changes, schema changes and ignored files needed for execution. A pull request description can omit an altered deployment permission or a generated worker configuration; the inventory must come from the proposed state.

Separate intended functionality from incidental rewrites. A new export endpoint does not automatically justify changing authentication helpers, broadening a storage policy or replacing the application framework. Ask the author to explain each changed boundary and remove unsupported changes through a reviewed revision. Preserve user-owned work outside the proposed change instead of resetting a shared worktree to produce a cleaner-looking diff.

Record AI assistance according to the repository's contribution policy without retaining confidential prompts unnecessarily. Note tools or externally obtained snippets when that information affects review, provenance or licensing questions. The author's responsibility does not move to the tool. Any unresolved attribution or license question needs an appropriate owner; this playbook provides an engineering procedure, not legal clearance.

3. Reconstruct the execution path and rank risks

Owner: independent reviewer with service maintainer. Output: boundary and risk register. Trace the request from authentication through authorization, record selection, job admission, worker execution, artifact storage and download. Include delayed and alternative paths such as scheduled cleanup, admin access and failed-job retries. Identify which changed function runs with which credentials and which authoritative record establishes completion.

For each boundary, name the defect that would change the release decision. In the fixture, examples include a cross-tenant query, a stale permission decision in the worker, an export that can be guessed by identifier and two retry attempts publishing different files under one operation identity. These examples are proposed test cases, not evidence that the candidate contains those defects.

Select specialist review where consequences require it. A data-model change may need migration review; a credential or policy change may need security review. Keep lower-risk changes proportionate rather than giving every file the same long checklist. A reviewer can accept a narrow scope or hold a dangerous boundary, but should not hide that boundary behind an overall score for the change.

4. Treat automated review as a source of questions

Owner: independent reviewer. Output: investigated findings with dispositions. Read assistant explanations and automated comments against the actual code and contract. Confirm locations, call paths and conditions before accepting a finding. Reproduce a suspected failure with synthetic data when possible. A convincing paragraph that refers to a deleted method is not evidence about the current candidate, and a tool's silence is not proof that a boundary is correct.

GitHub's Copilot Agents application card describes AI review as supplementary feedback and explains that generated responses require review and validation. Use that limitation when designing the workflow: ask the tool for likely gaps, then have the accountable reviewer establish which gaps are real. Do not claim that two models agreeing on an answer provide independent execution evidence.

Keep the assistant's working environment bounded. Repository files, issue text and package instructions can contain untrusted directions. A review aid should not receive deployment secrets, send real data to arbitrary services or execute a suggested fix against production. If another AI pass is useful, give it redacted evidence and a specific question. It can draft an explanation or a test proposal; it cannot approve its own consequential change.

5. Verify new dependencies and build-time execution

Owner: dependency maintainer with security reviewer. Output: dependency acceptance record. Inspect every added direct and transitive package, registry source, exact version and reason for inclusion. Verify the named project independently of the assistant's suggested URL. Check whether the existing approved library or platform feature already serves the requirement. A package's existence and popularity do not establish that it is the intended implementation or suitable for the production environment.

Inspect lifecycle scripts, native binaries, downloaded artifacts and credential access before installing the candidate in an isolated environment. Record the license review owner, supported runtime and unresolved provenance limits. Security scans inform the decision but do not verify the meaning of the package's API or cover every behavior in a bundled binary. Hold an unverified publisher or unexpected installation step rather than repeatedly installing until the demo works.

For an npm project, npm ci uses the lockfile and fails when dependency declarations do not match it, rather than updating the lockfile to fit. Use the reviewed installation settings in a disposable environment. A reproducible dependency tree is one evidence layer; it does not by itself prove that the packages are trustworthy or that the resulting application behaves correctly.

6. Test access through the request, worker and artifact

Owner: security reviewer with independent reviewer. Output: permission-boundary results. Construct synthetic identities with allowed, denied, expired and withdrawn access. Test two tenants whose records and artifacts are distinguishable. Attempt the operation through every exposed route, including direct retrieval and worker execution. Check both visible responses and the authoritative side effects: a denied response can still leave an unauthorized artifact behind.

OWASP's Authorization Cheat Sheet recommends permission validation on every request and a default-deny approach. Apply that guidance to the actual application policy. Do not substitute a successful login or an unguessable object identifier for authorization to the requested record. A worker acting after admission still needs the policy-defined execution authority; a download request needs the applicable artifact access check.

Review field selection and diagnostic exposure as well as object access. An allowed export may include a forbidden internal note or disclose another tenant's identifier through an error. Test the minimum permitted fields, retained artifacts and cleanup behavior. Record the exact revision, actor and policy behind each result. Missing fixtures for a changed access boundary are a hold, not evidence that the change leaves access unaffected.

These are parallel review concerns, not runtime services. Arrows bind evidence to the synthetic candidate and reviewer decision, not automatic approval. Missing required evidence prevents acceptance of the affected scope.

7. Build tests with independent expected results

Owner: independent reviewer with test maintainer. Output: inspected acceptance fixtures. Verify what each test executes, what it asserts and how its expected result was obtained. For the fixture, specify permitted records and fields before invoking the export. Inspect generated tests for assertions that merely compare output with the same implementation helper or accept any successful status. Such tests can run cleanly while reproducing the candidate's mistake.

Demonstrate that a consequential assertion can reject a known bad result in an isolated copy or fixture. For example, deliberately supply the other tenant's record to the output comparison and prove that the expected-data assertion fails. Restore the fixture and prove the correct result passes. This is a test-quality rehearsal, not permission to weaken production authorization or modify a shared branch to manufacture a failure.

Combine focused tests with the product boundary. Unit tests can check CSV encoding and identifier mapping, while a scoped integration fixture verifies request admission, worker behavior, storage and retrieval together. Track skipped tests, disabled assertions and unsupported environments explicitly. Coverage percentage cannot establish whether the permissions assertion exists or whether a mocked receiver concealed the external behavior introduced by the diff.

8. Exercise interruption, retries and unknown outcomes

Owner: service maintainer with independent reviewer. Output: failure and operation-settlement evidence. Interrupt the synthetic export at different stages: before job admission, after admission, during artifact creation and after storage succeeds but before the response reaches the caller. Record the operation identity and what the system can authoritatively establish at each point. A timeout after submission must not be automatically interpreted as no effect.

Retry the same permitted operation within the product's defined policy and inspect the resulting artifacts and job records. Decide whether the existing result is reused, a superseded result is retained or a new request is required. Do not let attempt identifiers replace the stable business-operation identity. Two successful worker attempts do not prove that one user operation produced only one intended result.

Test cancellation, authorization withdrawal and cleanup while work is pending. State which behavior is contractually required and which is only an operator convenience. Verify that a paused worker cannot later publish an artifact through stale authority. Keep unknown storage or notification outcomes in an owned investigation state until readback resolves them; do not report recovery solely because the worker process restarted.

9. Reproduce the candidate under the target execution constraints

Owner: release operator with change author. Output: build and environment record. Build from the pinned candidate and reviewed dependency settings, using synthetic configuration with production credentials absent. Record runtime versions, relevant flags, architecture and external adapters. Check that the candidate does not rely on an uncommitted local file, an implicit developer credential or a different package release fetched during the demonstration.

Exercise the changed resource bounds under the agreed workload. For an export worker, inspect memory, temporary storage, query pressure, concurrent jobs and cleanup on failure. Define the measured population and observation window. A small happy-path CSV cannot support a claim about large or concurrent exports, but the review need not invent a load forecast that the product owner has not supplied.

Check deploy-time configuration and permissions as part of the artifact's operating contract. If production uses a stricter storage policy or a different runtime, identify which tests cover that difference and which remain unverified. A green build in a permissive local environment does not settle the target permission question. Keep environmental gaps in the acceptance record with their owners and required evidence.

10. Rehearse rollback and retained-state compatibility

Owner: release operator with service maintainer. Output: recovery rehearsal and limits. Run the candidate, admit a synthetic export and then rehearse returning to the previous application revision. Identify what survives the code reversal: new job messages, schema changes, stored files, changed configuration and partially completed work. Check whether the old worker can interpret candidate-created messages and whether it would repeat an already completed effect.

If rollback requires a write pause, queue drain, schema compatibility layer or forward repair, state that dependency before approval. Test one candidate-created artifact and one uncertain operation through the recovery path. A flag that stops new export requests does not remove files already stored or revoke links already issued. Decide how those objects remain accessible, are withdrawn or enter manual review under the agreed product policy.

Record the point beyond which automatic reversal is unsafe. Some data changes cannot be undone by redeploying the old binary. The owner may accept a compatible incremental change, choose forward repair or hold the release until recovery is designed. Do not call the change reversible based solely on the ability to select a previous build in the deployment console.

11. Bind approval and checks to the actual release candidate

Owner: engineering lead with release operator. Output: accepted-scope or hold record. Collect the behavior contract, complete diff, dependency decisions, test results, access checks and recovery evidence. Record unresolved issues with consequence, owner and disposition. Identify the exact source revision and release artifact. If an assistant fixes a comment after review, inspect the new diff and rerun affected evidence before carrying approval forward.

GitHub's protected-branch documentation describes required reviews, status checks and settings for stale approvals. Verify the repository's actual rules, bypass permissions and check sources rather than assuming that a green merge button proves the intended policy is enforced. Feature availability depends on the account and repository setup. This playbook does not require buying a plan or changing live repository settings.

Use the following proposed acceptance criteria for the synthetic export change. The accountable owner must adapt them to the approved risk and scope; they are not a security certification.

  • The reviewed base, candidate, dependency tree and artifact are identified, with unexplained changes held.
  • Expected behavior comes from the agreed contract and synthetic inputs, not the candidate's own answers.
  • Allowed and denied tenant, worker and artifact paths have inspected results and side-effect readback.
  • Every new dependency has an identity, purpose and execution-boundary decision.
  • Required interruption, retry, cancellation and recovery cases have evidence or an explicit hold.
  • Approval is current for the candidate, and known gaps are visible to the release owner.

12. Verify controlled release evidence and keep accountability

Owner: release operator with engineering lead. Output: release readback and owned follow-up. After separately authorized deployment, verify that the deployed artifact matches the accepted identity. Check the changed operation in the approved exposure scope and observe the relevant errors, denied access, pending work and retained outputs. Use synthetic or explicitly authorized test records. Do not turn an engineering verification into an unsolicited real export or notification.

Keep hold criteria tied to the changed operation. Unexpected cross-tenant access, duplicate artifacts, incomplete observations or incompatible recovery state can require a pause even when the aggregate service dashboard is green. Stop new admission where authorized, preserve evidence and settle already admitted work. Record the difference between containing further exposure and completing recovery for existing operations.

For deeper package review, use Verify Dependencies Before Running AI-Generated Code. For retained-state reversal questions, read the feature-off recovery discussion. Ampity's backend systems and APIs and technology stack evaluation services can support the relevant review. Reading and downloading remain anonymous; asking us to contact you is optional.

Start the next code review with the pinned diff and one independently specified failure fixture. Assign a reviewer to the changed boundary that has no evidence yet, before asking an assistant to produce another summary of the candidate.