Skip to content
All case studies

LLM Automation + Developer Productivity

LLM-Based Pull Request Review Assistant

Built a review assistant that reads pull requests and leaves the comments a careful reviewer would leave, so the mechanical problems get caught before a person spends attention on them.

Mechanical review findings surfaced before human review begins

Review feedback made consistent regardless of reviewer load

Human attention redirected toward design judgment

Findings delivered inline, in the tool reviewers already use

The problem

Code review is the bottleneck nobody wants to name. Reviews queue behind whoever holds the context, quality varies with how much the reviewer has already read that day, and a large share of comments are things a careful reader catches mechanically. The valuable part of review, which is the judgment about whether the change is the right change at all, gets crowded out by the part that does not need judgment.

What I built

Built a workflow that pulls the change along with enough surrounding context to reason about it, then posts findings inline where a reviewer is already looking. The scope is deliberately narrow. It handles the mechanical layer and stays out of design judgment, because a tool that argues confidently about architecture gets muted within a week and then nobody gets the mechanical findings either.

Technical approach

  • The change is read with its surrounding context rather than as a bare diff, since a diff on its own cannot tell you whether a removed check was moved or lost
  • Findings post inline at the relevant line, because review that arrives as a separate report does not get read
  • Scope limited to what can be argued concretely, so the assistant does not spend its credibility on opinions
  • Tuned toward precision over coverage, because a reviewer who hits three false findings in a row stops reading the fourth
  • Runs inside the existing review workflow rather than asking developers to go somewhere new to get it

Visuals

Where the assistant sits in the existing review flow
Example finding, inline at the line it refers to
The precision and coverage trade-off, and where it was set
What the assistant is scoped out of, and why