Code Review Evidence When AI Wrote the Change
Review AI-written changes against independently defined behavior, failure paths and the actual tested revision, rather than relying on generated tests alone.
Review the behavior, not the author's identity
An AI-written change needs the same clear behavioral contract as any other change. Define what the user should experience, which state may change and which effects must not happen. Verify those conditions independently of the generated implementation and attach the evidence to the actual revision under review. AI authorship alone establishes neither safety nor a defect.
The particular risk is self-confirmation. An assistant can implement an assumption and generate tests that repeat it. Green tests then establish consistency between two pieces of code without proving either matches the requirement. A reviewer needs an independent source for expected behavior: an agreed product rule, existing contract, explicit acceptance example or system invariant.
This article uses a hypothetical document-download feature. Anyone may download a resource anonymously. A reader can separately request a reply by providing an email and context. The example is a review exercise, not a claim that a particular implementation or customer system has passed these checks. It makes the distinction between visible success and unwanted side effects concrete.
Write the acceptance rule before reading the test names
Start with plain language: downloading a document must not require contact information, and an anonymous download must not create a contact request. A reader asking for a reply should see a truthful confirmation of that request's delivery state. The document itself should remain available if the optional contact path fails.
Translate those rules into observable conditions. Can the reader proceed with both contact fields empty? Does the browser receive a valid file with a meaningful filename? Does the server receive a contact request only after an explicit choice? What does the interface say if that request is saved for later delivery instead of immediately confirmed?
Do not derive the expected result from the component's internal branch names. A function named downloadWithoutContact could still invoke a shared handler that writes to a CRM. The independent requirement concerns effects, not naming. Trace the path from user action to the systems capable of causing those effects.
The Google engineering guidance for code review asks reviewers to inspect functionality and whether tests would fail for broken behavior. Apply that question directly: which test would fail if the download started requiring an email again? If there is no such assertion, the most important requirement is not yet protected.
Make the tests challenge the implementation
A useful test crosses a meaningful boundary. For the download example, a UI test can inspect the received file and observe whether a contact endpoint was called. A server test can verify that missing reply intent is not transformed into a contact request. A delivery test can examine whether the receiving system contains the intended record after an explicit request.
These are different layers of evidence. A mocked server response cannot establish that a real CRM record exists. A CRM readback cannot establish that mobile users can reach the button. Label what each check actually proves rather than combining several partial checks into an unconditional success claim.
Challenge the test itself with a controlled defect in a disposable review setup, where appropriate. For example, temporarily make the download handler submit a contact request and confirm the no-contact assertion fails. Do not perform this experiment against production or actual recipients. The purpose is to establish that the assertion notices the prohibited effect, not to create an unwanted effect in a live system.
Manual review still matters. A test can select the wrong button or accept a generic response unrelated to the user's action. Inspect its setup, assertions and isolation. Keep fixtures independent of the production implementation's convenience helpers when sharing those helpers would reproduce the very mistake being tested.
Use a behavior evidence matrix
The following proposed matrix keeps outcomes and evidence separate. Replace the hypothetical scenarios with the actual product rules. It is not a universal test suite and does not substitute for security, accessibility or domain-specific review.
| Scenario | Required outcome | Evidence that addresses the requirement | | --- | --- | --- | | Reader downloads anonymously | File available without a reply request | File inspection plus absence of contact submission | | Reader asks for a reply | Explicit intent and provided context preserved | Request payload inspection and receiving-system readback | | Optional contact delivery fails | Download stays available; status is truthful | Injected failure and observed interface state | | Reader clicks twice | Effects follow the defined duplicate policy | Concurrent or repeated-action test at the responsible boundary | | Page loads on mobile | Action and explanation are usable | Visual inspection and interaction at supported widths | | A new revision changes the handler | Old evidence is not reused as current proof | Revision-bound test result and fresh review of the changed path |
['Reader downloads anonymously', 'File available without a reply request', 'File inspection plus absence of contact submission'], ['Reader asks for a reply', 'Explicit intent and provided context preserved', 'Request payload inspection and receiving-system readback'], ['Optional contact delivery fails', 'Download stays available; status is truthful', 'Injected failure and observed interface state'], ['Reader clicks twice', 'Effects follow the defined duplicate policy', 'Concurrent or repeated-action test at the responsible boundary'], ['Page loads on mobile', 'Action and explanation are usable', 'Visual inspection and interaction at supported widths'], ['A new revision changes the handler', 'Old evidence is not reused as current proof', 'Revision-bound test result and fresh review of the changed path'], ].map(([scenario, outcome, evidence]) => ( ))}
A table entry should lead to an executable check or an explicit manual observation. “Looks fine” is not enough for a prohibited server write. Conversely, a precise assertion about endpoint calls does not prove that the written copy is understandable. Choose the evidence that answers the question instead of forcing every requirement into one kind of test.
Inspect failure paths and intermediate states
Successful requests are usually the easiest path to demonstrate. Ask what happens after a timeout, an invalid file response, a browser interruption or a backend error. If the server accepted a contact request but the response was lost, blindly repeating it may create duplicates. The design needs a recovery policy appropriate to that effect, not just a spinner and another click.
Inspect user-facing status independently of internal status labels. A locally queued request is not a confirmed CRM delivery. A returned URL is not a verified PDF. A file with the right extension can still contain an HTML error page. Make the confirmation match the observed state and give the reader a useful next step when the outcome is uncertain.
State transitions also reveal accidental coupling. Disabling the download button because the optional email field is invalid makes an optional field effectively mandatory. Clearing a form before durable acceptance can lose the reader's context. Closing the dialog on every resolved promise can hide an application-level rejection carried inside a successful HTTP response.
Test the order of events where it matters. Suppose download succeeds but contact delivery fails. The reader should not be told that everything failed, nor should the interface imply a person will respond when there is no confirmed route for that request. Independent outcomes deserve separate status and recovery handling.
Bind evidence to the code that will actually run
Record the tested revision, environment, relevant configuration and dependencies. A screenshot from before the latest change cannot prove the latest behavior. Neither can a successful CI run for a different commit. If a fix modifies the event handler, rerun the checks that traverse that handler, including its failure and no-contact branches.
Inspect feature flags, runtime host checks and deployed API configuration. The same frontend source can behave differently if one environment points at a live contact endpoint and another uses a fixture. A staging pass should identify that boundary rather than present simulated delivery as real delivery. Production-only tracking also needs an environment-specific check, not an assumption based on a build passing.
Avoid making evidence collection itself a privacy problem. Use controlled fixtures, redact unnecessary personal information and keep secrets out of screenshots or logs. If a live canary is needed, obtain the appropriate authority and use a clearly identified test record. Local development permission does not automatically authorize contacting someone.
For high-risk behavior, review the produced artifact as well as the source. Confirm that the configuration and routes in the release candidate match the intended environment. The relevant result is what users receive, not what the developer intended the build to contain.
Use AI review as assistance, not independent approval
An AI reviewer can help locate changed branches, suggest missing cases or explain unfamiliar code. Its conclusions still need verification. The GitHub responsible-use guidance for Copilot agents describes missed problems, false positives and potentially incorrect suggestions. Treat a generated review comment as a hypothesis to investigate, not as proof of correctness.
Asking a second assistant to approve the first assistant's implementation does not by itself create independence. Both may inherit the same incomplete prompt, fixture or misunderstood product rule. Independence comes from grounding expectations outside the proposed implementation and checking actual outcomes at the relevant boundary.
Use automated review where it reduces search effort, then retain accountable human judgment for product meaning and consequential effects. A tool can suggest that email validation should be stricter. The reviewer must still notice whether changing that validation accidentally prevents anonymous downloads. Correct syntax can implement the wrong policy.
Do not convert every generated warning into a blocking requirement. Reproduce the concern, determine its scope and explain the decision. Unnecessary fixes can introduce new behavior or distract from the original change. Equally, a confident explanation without a reproducer should not dismiss a demonstrated prohibited effect.
Close the review with an honest scope statement
A useful review conclusion states what was checked and what remains unknown. For the hypothetical feature: anonymous downloading was exercised at specified widths, the file was inspected, no contact submission occurred on that path, and explicit reply requests were verified only against the fixture backend. That last boundary matters. Real delivery remains unproven until the receiving system is inspected.
If a gap is consequential, hold the change or narrow the release until the gap is resolved. Do not bury an untested write-recovery path under a list of unrelated green checks. If the gap is an acceptable limitation, give it an owner and a reason. Approval should communicate a decision, not merely a mood.
Preserve the important assertions as regression tests after review. The next change may be human-written, generated or mixed, but the user's contract stays the same. For dependency-specific checks, read verifying dependencies in AI-generated code. For implementation support, bring a concrete behavior and its evidence to product development, rather than asking for a generic AI safety badge.