README: DeepSWE benchmark results - #1511
Merged
Merged
Conversation
The Benchmarks section now carries the ten-harness DeepSWE comparison: the chart, a table of solved, cost per task, cost per solved issue and time, the limits of a one-seed run, and the V4.1 Flash and Kimi K3 runs. The per-harness numbers move into docs/benchmarks/deepswe. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit d8f8d3d)
The README keeps the chart and one paragraph; the table, the later V4.1 Flash and Kimi K3 runs, the setup and the limits live in docs/benchmarks/deepswe. The chart's footer now says only "same model". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit 9bdf25e)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit 7c92173)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit b1ac028)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE
The installer printed a three-line telemetry notice after the receipt. It prints none now; the binary's full notice still arrives before the first session's events are sent. The local install marker stays. The test asserts the installer prints no notice, and docs/TELEMETRY.md loses the installer's block. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit 53e9fd9)
Nothing the installer prints mentions telemetry now; the marker is still written when it can be. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE (cherry picked from commit df57499)
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE
AbirAbbas
added a commit
that referenced
this pull request
Sep 25, 2026
AbirAbbas
pushed a commit
that referenced
this pull request
Sep 25, 2026
…1436) A task's crew is picked per task by internal/crewroute (class, price on every connected route, quality minus lambda times cost) instead of a stored preset. /crew holds only what is allowed, what is pinned and a daily cap; --best and --cheap move one task. Old profiles migrate once. Route health, a per-seat fallback ladder, per-task and daily spend limits, /redo stronger, remote protocol 18. Added in review: a pin keeps its thinking level, auto in the reflex or small-work row reads its default and migrates, -yes-spend passes the daily cap, the cap refusal no longer offers --cheap, zero crew money is not drawn, and the manual describes the one seats row; dev merged in over #1429, #1485, #1494 and #1511. Co-authored-by: Santosh kumar <29346072+santoshkumarradha@users.noreply.github.com> Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
AbirAbbas
added a commit
that referenced
this pull request
Sep 25, 2026
…side senior-dev Brings in #1494 (team delegation), #1511 and #1436 (per-task worker, planner and checker routing). A senior-dev run keeps its own road beside them: it works alone in the folder it holds, so a via proposal starts one program run and does not go through the team split; it keeps the models it was asked for, or the pinned or profile worker, and #1436's per-task routing and $5 task cap apply to codeaf's own tasks, while senior-dev keeps its finite per-run ceiling. Both sides had claimed remote protocol version 18; the combined wire is version 19. Program badges fit #1494's grouped side column, the manuals describe the merged model choice, and the prompt-size ledger records the combined fixed prefix (57,218 bytes). dev's unbounded-launch test now holds its six workers at a barrier instead of guessing with a sleep, which failed on a loaded machine. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
#benchmarks./senior-dev… First of ten harnesses on DeepSWE".assets/readme/benchmark-deepswe.webp), one paragraph, and a link todocs/benchmarks/deepswe/.docs/benchmarks/deepswe/: the ten-harness table, the V4.1 Flash and Kimi K3 runs, setup and limits (README.md), and one row per harness (arms.csv), from the DeepSWE comparison run of 2026-09-12.test/installer-telemetry.shpasses, 28 of 28.Cherry-picked from
zeropoint95/senior-dev-harness(#1488) without its code.🤖 Generated with Claude Code
https://claude.ai/code/session_01V7ShhY74oyWjYGB3SougdE