An autonomous GitHub pull request reviewer built on LangGraph and TypeScript.
Structured findings · inline comments · sticky summary · cost per run.
PR Review Agent connects to your GitHub repositories and reviews pull requests the way a senior engineer would — catching real bugs, security issues, and bad patterns, not generating noise.
On every review it:
- Loads the PR context and diff, respects repo rules from
AGENTS.md/CLAUDE.md - Runs a multi-unit LLM review (fan-out across large diffs) with a dedicated verifier pass to drop speculative findings
- Merges CI check annotations on changed lines as first-class findings
- Posts up to
review.maxInlineCommentsinline comments and one sticky summary comment betweenpr-agent:reviewmarkers - Deduplicates findings on re-runs — pushes and redeliveries never create duplicate threads
- Refuses to publish when the head SHA has moved mid-run
The system has a transport-agnostic core (src/core/) shared by three shells:
| Shell | Trigger |
|---|---|
| CLI | pnpm dev review owner/repo#N |
| Webhook server | GitHub App event → POST /webhooks/github |
| Eval runner | Synthetic fixture JSON for offline quality measurement |
All shells call the same runCommand() → review graph pipeline. The core never knows which shell invoked it.
Review graph stages:
fetchContext → prepareDiff → reviewUnit (fan-out)
→ mergeFindings → mergeCiFindings → groundFindings
→ filterFindings → verifyFindings → summarize
Each finding is grounded against the actual diff line index before publishing. Ungrounded findings (model hallucinated a line number) are dropped, not posted.
Prerequisites: Node ≥ 24, pnpm
git clone https://github.com/your-org/pr-review-agent.git
cd pr-review-agent
pnpm install
cp .env.example .envSet the required environment variables in .env:
# Required for all modes
OPENAI_API_KEY=sk-...
# CLI with a personal token
GITHUB_TOKEN=ghp_...
# CLI or server with a GitHub App
GITHUB_APP_ID=123456
GITHUB_APP_PRIVATE_KEY="-----BEGIN RSA PRIVATE KEY-----\n..."
GITHUB_APP_INSTALLATION_ID=12345678 # CLI only
# Server mode
GITHUB_WEBHOOK_SECRET=your-webhook-secret# Dry run — review to stdout, nothing posted to GitHub
pnpm dev review owner/repo#12 --dry-run
# Post the review as inline comments + sticky summary
pnpm dev review owner/repo#12
# Generate a PR description (title, type, walkthrough)
pnpm dev describe owner/repo#12 --dry-run
# Inspect the effective merged configuration
pnpm dev config show
# Fetch and print raw PR data (smoke test)
pnpm smoke:github owner/repo#12Global flags: --config <file>, --log-level debug|info|warn|error, --json, --record-dir <dir> (saves run records).
# Development (with hot reload)
pnpm dev:server
# Production
pnpm build && pnpm start
# Docker
docker build -t pr-review-agent .
docker run -p 3000:3000 --env-file .env pr-review-agentThe server handles:
| Event | Action |
|---|---|
pull_request opened / synchronize / reopened / ready_for_review |
Auto-review (if server.autoReviewOnOpen: true) |
issue_comment /review or /describe |
Triggered review by a collaborator |
GET /healthz |
Liveness check |
See docs/WEBHOOK.md for GitHub App permissions and smee.io tunnel setup for local development.
Collaborators only. Bot authors and drive-by commenters are rejected.
/review — trigger a fresh review on the current head
/describe — regenerate the PR description
Drop a .pr-agent.yaml file on the default branch of any repository. Policy lives in the repo, not in the bot deployment.
review:
maxInlineComments: 10 # cap on posted inline comments (overflow in summary)
enableVerifier: true # second-pass keep/drop per finding (recommended)
minConfidence: 0.6 # drop findings below this threshold
severityThreshold: low # minimum severity to post
maxConcurrency: 4 # parallel review units and verifier batches
diff:
maxFiles: 30 # files included in the diff
contextLines: 4 # extra lines around each hunk
ignore:
globs:
- "pnpm-lock.yaml"
- "*.generated.ts"
models:
reviewer:
provider: openai
model: gpt-4o
verifier:
provider: openai
model: gpt-4o-mini
summarizer:
provider: openai
model: gpt-4o-mini
budgets:
perRunUsd: 0.50 # hard stop (no silent fallback)
perRunTokens: 200000Configuration layers (last wins): built-in defaults → .pr-agent.yaml on default branch → CLI --config flag. Invalid layers are dropped entirely and logged — they never silently corrupt the config.
The eval suite is the primary quality gate. It runs the real review graph against synthetic fixture PRs and scores precision, recall, and F1 against labeled line ranges.
pnpm evalExample output:
| Metric | Value |
|---|---|
| Cases | 12 synthetic PRs (bugs, security, clean, adversarial) |
| Precision | Reported per run |
| Recall | Reported per run |
| Cost | Tokens and USD per case in the JSON report |
Record a baseline after a clean run, then compare future runs against it to catch regressions before they reach production. See evals/baselines/README.md.
Eval fixtures cover: missing await, SQL injection, hardcoded secret, breaking API change, off-by-one error, weak crypto, open redirect, empty catch, prompt injection, and clean PRs.
Runs typecheck, lint, format check, and all tests:
pnpm checkIndividual commands:
pnpm typecheck
pnpm lint
pnpm test
pnpm test:coverageThe server is a single Node process. Deploy anywhere that runs Docker or Node 24.
Fly.io (example — adjust app name in fly.toml):
fly secrets set OPENAI_API_KEY=... GITHUB_APP_ID=... GITHUB_APP_PRIVATE_KEY=... GITHUB_WEBHOOK_SECRET=...
fly deployDocker:
docker build -t pr-review-agent .
docker run -p 3000:3000 --env-file .env pr-review-agentThe /healthz endpoint returns { "ok": true } and is used as the liveness check.
Note on scaling: In-memory dedupe and per-PR run locks are per-process. Multiple replicas need sticky routing or external state. For a single-instance deployment this is not an issue.
| Area | State |
|---|---|
| Core, CLI, diff engine | Done |
| Review graph (fan-out, grounding, verifier) | Done |
| CI findings from check annotations | Done |
| Webhook server + skip rules + slash commands | Done |
| Evals (precision/recall, baseline comparison) | Done |
| Dockerfile + Fly.io deploy | Done |
| Queue / Postgres / Redis | Out of scope |
| Pinecone retrieval, specialist fleet, agent tools | Out of scope |
src/
core/ All product logic (config, diff, context, llm, graph, github, commands)
server/ Hono webhook server, supervisor, delivery cache
cli/ Commander entry point
evals/ Eval runner, scoring, fixture cases, reports
demo/ Seed instructions for a public demo repository
docs/ MVP-PLAN.md, WEBHOOK.md, IMPLEMENTATION.md
tests/ Vitest unit, integration, and graph tests
