Recover the contract before reading the explanation
An AI-generated PR can arrive with a polished explanation and plausible tests. Begin with the issue or a concrete user scenario, then write down what success requires in your own words. That gives you a reference independent of the proposed implementation.
Suppose the request is “let repository members reopen a saved review.” The contract includes which members may read it, what happens if their access changes, and whether reopening restores an existing result or starts a new analysis. “Adds a history endpoint” answers only a small part of that request.
Keep this note short: one success case, one rejection case, and the existing behavior that must survive.
Inspect the boundaries where assumptions meet reality
Follow one request from its entry point to its final effect. Pay particular attention to transitions between systems or owners:
- Identity → authorization: being signed in must not stand in for permission to this repository or record.
- Input → stored state: check malformed values, missing records, and who controls each identifier.
- Network → application state: inspect timeouts, partial responses, retries, and duplicate effects.
- Server → UI: verify that empty, loading, and failed results are distinguishable.
Check newly introduced APIs against their actual definitions and versions. If the explanation says a helper validates access, open that helper and its caller. Names and comments are clues; the executed path is the evidence.
Choose a check that could prove the change wrong
Tests written alongside the code may inherit its assumptions. Ask what should happen when an input is close to valid but forbidden: a different repository, an expired session, or a repeated request after a timeout.
For saved reviews, one useful check is whether a user who can access repository A can read a saved run belonging to repository B. Assert the externally visible result and that private content is absent. A test that merely checks whether the implementation called an authorization helper can miss an incorrect decision.
Also inspect the successful path. A rejection test is not enough if legitimate members can no longer open their own reviews.
Make the decision from evidence
Before approving, record the behavior you checked, the relevant source, and the remaining uncertainty. When evidence is missing, name the missing check precisely so the author can resolve it.
AI assistance can help organize a review, but the output still needs inspection. GitHub’s documentation on Copilot code review describes limitations of AI review; its pull request review reference explains the review states used to communicate a decision.
Try Diff it’s demo to inspect behavior explanations and their evidence. Use those explanations as review leads, then check the code and relevant execution results yourself.