Skip to content

Repository files navigation

autonomous-dev-loop

Tests Evals

An AI dev loop that isn't allowed to wreck your repo.

Label an issue → an LLM writes the PR → a second LLM reviews it against real test/lint results → an auto-fix agent addresses the findings → you merge. Runs entirely on GitHub Actions; Groq (openai/gpt-oss-120b) by default, Anthropic optional.

Asked to fix one review finding, the auto-fix agent replaced a 690-line, 26-test suite with an 18-line stub that couldn't run (ADR-0009). Each boundary below links to the ADR recording the incident or exposure that forced it.

Dogfooded The loop wrote 13 merged PRs of its own (11 features, 2 bug fixes — error taxonomy, bounded retry, checkpoint resume, provider fallback), each merged after human review. Then it was locked out of its own code: scripts/, prompts/ and config/ are now on the write denylist (ADR-0021).
Tested 800+ tests on the built-in node:test runner, zero test dependencies, smoke tests wired to the real prompts and config
Measured Offline evals run each LLM stage against a labelled dataset: verdict accuracy, per-class precision/recall/F1, error rate, consistency across repeats, latency, tokens — gated by thresholds, latest results on the scorecard (ADR-0027)
Decided in writing 27 ADRs, each with context, rejected alternatives and trade-offs

Bounded autonomy, not "fully autonomous"

Prompts are advice; an LLM can ignore them. So the loop separates what it asks the model from what it enforces in code:

Enforced in code / CI Where
The review verdict is forced to REQUEST_CHANGES when any declared check (here: the test suite and a node --check syntax pass) fails on the PR head. A check that times out or crashes, or evidence that is missing or stale, withholds approval without triggering auto-fix (ADR-0026) ADR-0024
PR code runs in a job holding no secrets (contents: read, no persisted credentials, credential-like env vars stripped) ADR-0024
Pipeline scripts, prompts and config always run from the default branch — a PR cannot rewrite the code that reviews it ADR-0023
Write denylist: the model cannot touch .github/, scripts/, prompts/, config/, lockfiles, .npmrc, or escape via symlinks ADR-0021
Max 6 files per run, no absolute paths, no .., 16 000 chars per file ADR-0003
Max 3 auto-fix attempts, then escalation to a human; per-PR concurrency ADR-0006, ADR-0020
Under-specified issues never reach generation (validation gate, ready-for-dev label) ADR-0001
Human merge — the loop never merges: no merge call exists in scripts/ or the workflows docs/mvp.md (human review before merge), .github/workflows/
Asked of the model (prompt guardrails) Where
Never shrink a test file, never mix ESM/CJS, never change an exported signature, never add an undeclared package, never rewrite > 30% of a file ADR-0009
Named defect checklist in review (read-only property writes, unauthorized imports, non-persistent refs), disclosure when the diff was truncated docs/code-generation.md

Moving more of the second table into the first is the open work — see the proposed static verification backstop (ADR-0019).

How it works

graph LR
    A[Issue created] --> B[Validator]
    B -->|invalid| Z[needs-refinement — no PR]
    B -->|valid| C[label: ready-for-dev]
    C --> D[Code Generation]
    D --> E[PR opened / push]
    E --> EV[Evidence job<br/>no secrets]
    EV -->|review-evidence.json| F[PR Review]
    F -->|APPROVE| G[Human merge gate]
    F -->|REQUEST_CHANGES<br/>forced on any failing check| H{Attempt ≤ 3?}
    H -->|Yes| I[Auto-Fix]
    I --> F
    H -->|No| J[Manual intervention requested]
Loading

Quick start

  1. Add GROQ_API_KEY (or ANTHROPIC_API_KEY) in Settings → Secrets and variables → Actions. AI_PR_TOKEN is recommended for PR/label/review writes.
  2. Open an issue. The validator scores it and applies ready-for-dev or needs-refinement; labels are created automatically.
  3. Watch the PR appear, get reviewed, and get fixed. Merge it yourself.

Full setup, permissions, per-stage model keys and the end-to-end test: docs/code-generation.md. Failure triage: docs/runbook.md.

Code map

  • Workflows: .github/workflows/ (orchestration only)
  • Entrypoints: scripts/*.mjs, modules: scripts/lib/*.mjs
  • Prompts: prompts/*.md (one file per prompt, loaded at runtime)
  • Evals: scripts/run_evals.mjs, suites in scripts/lib/eval_suites.mjs, datasets in evals/datasets/
  • MVP definition: docs/mvp.md

Iterative Review Loop

Once a PR is opened, the automation continues:

  1. PR Review (.github/workflows/pr-review.yml) — triggered on every push to the PR branch. Posts or updates a review comment and submits an APPROVE or REQUEST_CHANGES verdict.
  2. Auto-Fix (.github/workflows/auto-fix-pr.yml) — triggered when a review requests changes. Reads the review feedback, generates targeted fixes using the LLM, and pushes them back to the PR branch — re-triggering the review.

The loop runs up to 3 auto-fix iterations per PR. After that, a comment is posted requesting manual intervention.

Fail-Fast Startup & Payload Validation

Automation entrypoints now validate critical runtime inputs before network calls:

  • required env vars are validated up-front with explicit errors,
  • required prompt files are validated as existing and non-empty at load time,
  • GitHub event payload fields are validated with explicit path-oriented messages (for example pull_request.number, issue.number, pull_request.head.ref / ref),
  • provider response parsing errors include concrete JSON paths (content[0].text, choices[0].message.content).

Observability

Every pipeline run produces two complementary outputs:

  • Structured JSON events — one JSON line per event written to stderr by each script (ts, run_id, stage, event, level, duration_ms, meta). Error-level events also emit ::error:: GitHub Actions annotations.
  • Run trace file — observability/traces/<GITHUB_RUN_ID>.json, written incrementally so it is always readable mid-run. Uploaded as the artifact run-trace-<GITHUB_RUN_ID> at the end of each workflow (if: always()).

Read a trace locally:

cat observability/traces/<run_id>.json | jq '[.spans[] | {stage, outcome, duration_ms}]'

All instrumentation goes through scripts/lib/observability.mjs. Schema reference and full event tables: docs/observability.md. Design rationale: ADR-0018.

Evals

Tests mock the LLM to prove the wiring; evals call the real model on a fixed, labelled dataset to measure the quality of its decisions — so a prompt or model change can be compared on the same inputs before merge.

Suite Stage Dataset Gate (exit 1 below)
validation Issue validation 15 issues — valid, and each blocker B1–B4 verdict accuracy ≥ 0.8 · invalid recall ≥ 0.8 · error rate ≤ 0.05
npm run eval -- --suite validation --repeats 3     # live, needs GROQ_API_KEY or ANTHROPIC_API_KEY
npm run eval -- --suite validation --repeats 3 --scorecard   # + record on the scorecard below
npm run eval -- --suite validation --replay evals/results/validation-<runId>.json   # re-score, no LLM call

Latest results

No live eval run recorded yet — see evals/SCORECARD.md for how to record one.

Each run prints a Markdown report (metrics, threshold failures, failing cases), writes the full replayable results to evals/results/, and appends a summary to evals/history.jsonl. In CI: Actions → Evals → Run workflow — the report lands on the run's summary page, results in the artifacts, and a run on main opens a chore(evals): scorecard update PR that updates the scorecard below and evals/SCORECARD.md once merged.

Adding a case, a scorer or a suite for another stage: docs/evals.md. Design rationale: ADR-0027.

Tests

The test suite uses the built-in node:test runner — no external dependencies.

node --test scripts/tests/*.test.mjs

Two layers of tests:

  • Unit tests — each module tested in isolation (config, output_writer, issue_validator, observability, etc.)
  • Smoke tests (smoke.test.mjs) — full pipelines with real config files and prompt templates, LLM mocked at the network boundary

CI: .github/workflows/test.yml runs the full suite on every push and PR. Guide: docs/testing.md.

About

Issue → PR → AI review → auto-fix, on GitHub Actions. The model proposes, code enforces (failing checks block approval, write denylist, 3-attempt cap), a human merges.

Topics

Resources

Contributing

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages