Repository navigation
Report coverage for incomplete legacy assigner runs - #239
Open
appleweiping wants to merge 2 commits into
Open
appleweiping wants to merge 2 commits into
appleweiping wants to merge 2 commits into
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The legacy analysis reads directories with
overall.json, so an interrupted run can leave useful logs without appearing in the summary.This adds a separate coverage command for the legacy assigner. It saves the actual task indices before running, counts unique terminal samples, and reports pending samples and infrastructure errors even when
overall.jsonis missing. A normal terminal model failure still counts as evaluated. Historical runs without a manifest keep an unknown denominator.The command does not calculate scores or modify existing score files. Retry entries and duplicate results do not inflate the sample count. Conflicting terminal records are reported instead of selecting the better result. Recorded attempt counts are explicitly log-entry counts because the old assigner does not log NOT_AVAILABLE callbacks. Changed task indices on resume require a new output directory.
Run it with:
I ran 14 tests on Windows with Python 3.12. They cover 8 completed samples out of 10, the public CLI, malformed logs, missing manifests, duplicate and conflicting results, and the real assigner initializer and resume with in-process clients. Ruff and diff checks pass. The added CI job targets Python 3.9 and 3.12; those remote jobs have not run yet.
This is limited to the legacy runner still documented in the README. It does not support AgentRL FC logs and does not change retry budgets or official benchmark scoring.