Skip to content

Report coverage for incomplete legacy assigner runs - #239

Open
appleweiping wants to merge 2 commits into
THUDM:mainfrom
appleweiping:feat/run-coverage-report
Open

appleweiping wants to merge 2 commits into
THUDM:mainfrom
appleweiping:feat/run-coverage-report

Conversation

@appleweiping

Copy link
Copy Markdown

The legacy analysis reads directories with overall.json, so an interrupted run can leave useful logs without appearing in the summary.

This adds a separate coverage command for the legacy assigner. It saves the actual task indices before running, counts unique terminal samples, and reports pending samples and infrastructure errors even when overall.json is missing. A normal terminal model failure still counts as evaluated. Historical runs without a manifest keep an unknown denominator.

The command does not calculate scores or modify existing score files. Retry entries and duplicate results do not inflate the sample count. Conflicting terminal records are reported instead of selecting the better result. Recorded attempt counts are explicitly log-entry counts because the old assigner does not log NOT_AVAILABLE callbacks. Changed task indices on resume require a new output directory.

Run it with:

python -m src.run_coverage --output outputs --save analysis/coverage.json

I ran 14 tests on Windows with Python 3.12. They cover 8 completed samples out of 10, the public CLI, malformed logs, missing manifests, duplicate and conflicting results, and the real assigner initializer and resume with in-process clients. Ruff and diff checks pass. The added CI job targets Python 3.9 and 3.12; those remote jobs have not run yet.

This is limited to the legacy runner still documented in the README. It does not support AgentRL FC logs and does not change retry budgets or official benchmark scoring.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant