Skip to content

fix(queue): report actual executed model in completion evidence - #704

Merged
ifuri-validator-agent[bot] merged 1 commit into
mainfrom
ticket/445-actual-runtime-model
Oct 8, 2026
Merged

ifuri-validator-agent[bot] merged 1 commit into
mainfrom
ticket/445-actual-runtime-model

Conversation

@tom-sapletta-com

Copy link
Copy Markdown
Contributor

Problem

A live Willman -> Koru -> SubLLM run using local GPT-OSS returned the expected answer but labeled its completion and run log as llm cursor/grok-4.6. The queue label used the requested/default model rather than the executed result metadata.

Change

Use the optional runner result.model before the requested/default label. Preserve legacy runners without model metadata and all return codes. Scope: ticket-445; intake paxlet-com/willman PLF-115 under user-authorized autonomous bug fixing.

Validation

  • 45 targeted tests passed, including ten new provenance/failure/fallback cases.
  • Managed governance PASS, zero warnings; source/test lint PASS.
  • Real minis integration run returned MINIS_KORU_LOCAL_OK, completed, and now reports llm local-development/gpt-oss:20b (4.35 seconds).
  • Protected exact-head OneDev checks and independent Validator approval required before merge.

This change updates execution evidence. Continuous fleet writer admission remains separate work; no fleet writer is enabled.

@ifuri-validator-agent ifuri-validator-agent Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Validator approval after policy checks for exact head 04a5494f267c5351cb6cfe26979f066ecce90b91.

Ticket: ticket-445
Correlation ID: local-semcod-koru-pr-704-ticket-445
Model: openai/cursor-auto
Reviewed diff chunks: 1
Advisory LLM verdict: APPROVE
Advisory summary: Reviewed all 1 diff chunk(s). The change correctly prioritizes the executed runner result's optional model metadata while preserving requested-model and default-model fallbacks. The regression tests cover runtime provenance, legacy results without model metadata, empty metadata, default fallback, and success/failure return codes. All protected required checks passed.
Advisory findings: none
The LLM output above is advisory and was not used as the approval trust root.
Semantic review prerequisite: not_required; policy 676cb4516bbfed2a000e40b9b1b6e4a430ecc761ec546aeb53d721a1905cfdd7.

Actual PR impact radar

Exact range: 552ed81026c9213d05029c9ce395d2e01fa0ca7a...04a5494f267c5351cb6cfe26979f066ecce90b91
Change digest: 35847f43bdfdc0a5d78a18c59a722822b77a29089d4d4f5653aadec63fceb432
Score: 68/100 (L), estimated 90 min, split recommended: true
Affected services/components: repository-wide/unclassified

Machine-readable radar JSONL and SVG
{"actual_change":{"additions":165,"base_sha":"552ed81026c9213d05029c9ce395d2e01fa0ca7a","binary_files":0,"categories":{"code":1,"configuration":1,"docs":2,"tests":1},"change_digest":"35847f43bdfdc0a5d78a18c59a722822b77a29089d4d4f5653aadec63fceb432","comparison":"552ed81026c9213d05029c9ce395d2e01fa0ca7a...04a5494f267c5351cb6cfe26979f066ecce90b91","deletions":1,"file_count":5,"files":["project/TICKETS.md","project/ticket-445/README.md","project/ticket-445/intent.json","src/koru/queue/runner.py","tests/test_queue_model_provenance.py"],"head_sha":"04a5494f267c5351cb6cfe26979f066ecce90b91","service_count":0,"services":[]},"assessment_mode":"observed-pr","axes":{"coupling":5,"delivery":3,"scope":5,"uncertainty":3,"validation":1},"complexity":"L","confidence":0.9,"diagnostics":["RADAR-ACCEPTANCE-MISSING","RADAR-BUDGET-EXCEEDED"],"estimate":{"budget_minutes":30,"minutes":90,"within_budget":false},"impact":{"components":["cursor","local-development","paxlet-com","project","provenance","requested","source","src/koru","tests"],"files":["cursor/grok-4.6","local-development/gpt-oss","paxlet-com/willman","project/TICKETS.md","project/ticket-445/README.md","project/ticket-445/intent.json","provenance/failure/fallback","requested/default","source/test","src/koru/queue/runner.py","tests/test_queue_model_provenance.py"],"public_interfaces":[],"runtime_dependencies":0},"schema":"subactor.ticket-radar/v1","score":68,"split":{"parts":[{"estimated_minutes":10,"name":"Implement cursor","scope":["cursor"]},{"estimated_minutes":10,"name":"Implement local-development","scope":["local-development"]},{"estimated_minutes":10,"name":"Implement paxlet-com","scope":["paxlet-com"]},{"estimated_minutes":10,"name":"Implement project","scope":["project"]},{"estimated_minutes":10,"name":"Implement provenance","scope":["provenance"]},{"estimated_minutes":15,"name":"Validate and project to trackers","scope":["tests","planfile","github/gitlab/jira projections"]}],"reason":"estimated_minutes_exceed_budget","recommended":true},"standards":[{"id":"wellmanifest/dsl","revision":"6c60fc4e0dd1f1bb74f46a7745e28019908d1203","version":"0.1.0-dev"},{"id":"wellmanifest/ticket-lifecycle","revision":"5bf581907a87b46a13a73e6c033d3abe4d9a306f","version":"0.1.0-dev"},{"id":"wellmanifest/git-lifecycle","revision":"7d77d4b7af57e69bc75c3a0290b3a4805c5c4438","version":"0.2.0-dev"},{"id":"wellmanifest/logs","revision":"48c284ef7a069055c0bcb6b900147ce5e65f8b43","version":"0.3.0"}],"ticket_ref":"ticket-445"}
<svg xmlns="http://www.w3.org/2000/svg" width="128" height="128" viewBox="0 0 128 128" role="img"><title>ticket-445: fix(queue): report actual executed model in completion evidence</title><rect width="128" height="128" rx="12" fill="#f8fafc"/><g stroke-width="1"><polygon points="64,55 72,61 69,71 59,71 56,61" fill="none" stroke="#d7dde5"/><polygon points="64,47 80,59 74,78 54,78 48,59" fill="none" stroke="#d7dde5"/><polygon points="64,38 89,56 79,85 49,85 39,56" fill="none" stroke="#d7dde5"/><polygon points="64,30 97,53 84,92 44,92 31,53" fill="none" stroke="#d7dde5"/><polygon points="64,21 105,51 89,99 39,99 23,51" fill="none" stroke="#d7dde5"/><line x1="64" y1="64" x2="64" y2="21" stroke="#aab4c0"/><line x1="64" y1="64" x2="105" y2="51" stroke="#aab4c0"/><line x1="64" y1="64" x2="89" y2="99" stroke="#aab4c0"/><line x1="64" y1="64" x2="39" y2="99" stroke="#aab4c0"/><line x1="64" y1="64" x2="23" y2="51" stroke="#aab4c0"/></g><polygon points="64,21 105,51 79,85 59,71 39,56" fill="#fb923c" fill-opacity="0.45" stroke="#c2410c" stroke-width="2"/><circle cx="64" cy="64" r="3" fill="#c2410c"/><g font-family="sans-serif" font-size="7" fill="#334155"><text x="64" y="11" text-anchor="middle">SCO</text><text x="114" y="48" text-anchor="middle">COU</text><text x="95" y="107" text-anchor="middle">UNC</text><text x="33" y="107" text-anchor="middle">VAL</text><text x="14" y="48" text-anchor="middle">DEL</text></g><text x="64" y="124" text-anchor="middle" font-family="sans-serif" font-size="8" fill="#0f172a">L · 90m</text></svg>
Merge will be attempted after this approval when explicitly authorized. ## Decision record (recomputable)
DECISION D-445-0939
TICKET ticket-445
HEAD_SHA 04a5494f267c5351cb6cfe26979f066ecce90b91
CORRELATION_ID local-semcod-koru-pr-704-ticket-445
ACTOR agent:ifuri-validator-agent[bot]
APPLIED_RULE P-CORE-015
INPUT author_login = "tom-sapletta-com"
INPUT observed_checks = ["standard packs / conformance=PASS","governance / enforce=PASS","governance / remote lifecycle=PASS","onedev/local-verify=PASS"]
INPUT required_checks = ["onedev/local-verify","standard packs / conformance"]
INPUT required_checks_source = "protected registry + GitHub applied rules (env/request)"
INPUT reviewer_login = "ifuri-validator-agent[bot]"
INPUT semantic_review_assessment = {"schema":"subactor.validator/semantic-review-assessment/v1","subject":{"repository":"semcod/koru","pull_request":704,"head_sha":"04a5494f267c5351cb6cfe26979f066ecce90b91","base_sha":"552ed81026c9213d05029c9ce395d2e01fa0ca7a","diff_sha256":"fa532f63a2602f11d7f4add2255e7652611df9abb6f00db9bb838c385927a424"},"policy":{"policy_schema":"subactor.validator/semantic-review-policy/v1","policy_version":1,"policy_sha256":"676cb4516bbfed2a000e40b9b1b6e4a430ecc761ec546aeb53d721a1905cfdd7","required":false,"critical_paths":[],"observed_paths":["project/TICKETS.md","project/ticket-445/README.md","project/ticket-445/intent.json","src/koru/queue/runner.py","tests/test_queue_model_provenance.py"]},"grounding":"full-diff-not-per-finding-proof","execution_authority":false,"status":"not_required","reason":null,"review_sha256":null,"unresolved":[]}
INPUT superseded_checks = ["smoke"]
INPUT ticket_radar_receipt = {"schema":"subactor.ticket-radar/v1","base_sha":"552ed81026c9213d05029c9ce395d2e01fa0ca7a","head_sha":"04a5494f267c5351cb6cfe26979f066ecce90b91","change_digest":"35847f43bdfdc0a5d78a18c59a722822b77a29089d4d4f5653aadec63fceb432","score":68,"complexity":"L","estimated_minutes":90,"split_recommended":true,"services":[],"authority":"ADVISORY","promotion":"FORBIDDEN"}
VERDICT APPROVE AUTHORITY DETERMINISTIC
REJECTED REQUEST_CHANGES BECAUSE NO_UNSAFE_CHANGE_REASON_FOUND
ADVISORY llm_verdict = "APPROVE" MODEL "openai/cursor-auto"
ASSERT VERDICT_AUTHORITY != "ADVISORY"

@ifuri-validator-agent ifuri-validator-agent Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Validator approval after policy checks for exact head 04a5494f267c5351cb6cfe26979f066ecce90b91.

Ticket: ticket-445
Correlation ID: local-semcod-koru-pr-704-ticket-445
Model: openai/cursor-auto
Reviewed diff chunks: 1
Advisory LLM verdict: APPROVE
Advisory summary: Reviewed all 1 diff chunk(s). The visible change correctly prioritizes the executed runner result model while preserving requested-model and default-model fallbacks. Tests cover runtime provenance, legacy absence, empty values, default fallback, and nonzero exit-code preservation. All protected required checks passed.
Advisory findings: none
The LLM output above is advisory and was not used as the approval trust root.
Semantic review prerequisite: not_required; policy 676cb4516bbfed2a000e40b9b1b6e4a430ecc761ec546aeb53d721a1905cfdd7.

Actual PR impact radar

Exact range: 552ed81026c9213d05029c9ce395d2e01fa0ca7a...04a5494f267c5351cb6cfe26979f066ecce90b91
Change digest: 35847f43bdfdc0a5d78a18c59a722822b77a29089d4d4f5653aadec63fceb432
Score: 68/100 (L), estimated 90 min, split recommended: true
Affected services/components: repository-wide/unclassified

Machine-readable radar JSONL and SVG
{"actual_change":{"additions":165,"base_sha":"552ed81026c9213d05029c9ce395d2e01fa0ca7a","binary_files":0,"categories":{"code":1,"configuration":1,"docs":2,"tests":1},"change_digest":"35847f43bdfdc0a5d78a18c59a722822b77a29089d4d4f5653aadec63fceb432","comparison":"552ed81026c9213d05029c9ce395d2e01fa0ca7a...04a5494f267c5351cb6cfe26979f066ecce90b91","deletions":1,"file_count":5,"files":["project/TICKETS.md","project/ticket-445/README.md","project/ticket-445/intent.json","src/koru/queue/runner.py","tests/test_queue_model_provenance.py"],"head_sha":"04a5494f267c5351cb6cfe26979f066ecce90b91","service_count":0,"services":[]},"assessment_mode":"observed-pr","axes":{"coupling":5,"delivery":3,"scope":5,"uncertainty":3,"validation":1},"complexity":"L","confidence":0.9,"diagnostics":["RADAR-ACCEPTANCE-MISSING","RADAR-BUDGET-EXCEEDED"],"estimate":{"budget_minutes":30,"minutes":90,"within_budget":false},"impact":{"components":["cursor","local-development","paxlet-com","project","provenance","requested","source","src/koru","tests"],"files":["cursor/grok-4.6","local-development/gpt-oss","paxlet-com/willman","project/TICKETS.md","project/ticket-445/README.md","project/ticket-445/intent.json","provenance/failure/fallback","requested/default","source/test","src/koru/queue/runner.py","tests/test_queue_model_provenance.py"],"public_interfaces":[],"runtime_dependencies":0},"schema":"subactor.ticket-radar/v1","score":68,"split":{"parts":[{"estimated_minutes":10,"name":"Implement cursor","scope":["cursor"]},{"estimated_minutes":10,"name":"Implement local-development","scope":["local-development"]},{"estimated_minutes":10,"name":"Implement paxlet-com","scope":["paxlet-com"]},{"estimated_minutes":10,"name":"Implement project","scope":["project"]},{"estimated_minutes":10,"name":"Implement provenance","scope":["provenance"]},{"estimated_minutes":15,"name":"Validate and project to trackers","scope":["tests","planfile","github/gitlab/jira projections"]}],"reason":"estimated_minutes_exceed_budget","recommended":true},"standards":[{"id":"wellmanifest/dsl","revision":"6c60fc4e0dd1f1bb74f46a7745e28019908d1203","version":"0.1.0-dev"},{"id":"wellmanifest/ticket-lifecycle","revision":"5bf581907a87b46a13a73e6c033d3abe4d9a306f","version":"0.1.0-dev"},{"id":"wellmanifest/git-lifecycle","revision":"7d77d4b7af57e69bc75c3a0290b3a4805c5c4438","version":"0.2.0-dev"},{"id":"wellmanifest/logs","revision":"48c284ef7a069055c0bcb6b900147ce5e65f8b43","version":"0.3.0"}],"ticket_ref":"ticket-445"}
<svg xmlns="http://www.w3.org/2000/svg" width="128" height="128" viewBox="0 0 128 128" role="img"><title>ticket-445: fix(queue): report actual executed model in completion evidence</title><rect width="128" height="128" rx="12" fill="#f8fafc"/><g stroke-width="1"><polygon points="64,55 72,61 69,71 59,71 56,61" fill="none" stroke="#d7dde5"/><polygon points="64,47 80,59 74,78 54,78 48,59" fill="none" stroke="#d7dde5"/><polygon points="64,38 89,56 79,85 49,85 39,56" fill="none" stroke="#d7dde5"/><polygon points="64,30 97,53 84,92 44,92 31,53" fill="none" stroke="#d7dde5"/><polygon points="64,21 105,51 89,99 39,99 23,51" fill="none" stroke="#d7dde5"/><line x1="64" y1="64" x2="64" y2="21" stroke="#aab4c0"/><line x1="64" y1="64" x2="105" y2="51" stroke="#aab4c0"/><line x1="64" y1="64" x2="89" y2="99" stroke="#aab4c0"/><line x1="64" y1="64" x2="39" y2="99" stroke="#aab4c0"/><line x1="64" y1="64" x2="23" y2="51" stroke="#aab4c0"/></g><polygon points="64,21 105,51 79,85 59,71 39,56" fill="#fb923c" fill-opacity="0.45" stroke="#c2410c" stroke-width="2"/><circle cx="64" cy="64" r="3" fill="#c2410c"/><g font-family="sans-serif" font-size="7" fill="#334155"><text x="64" y="11" text-anchor="middle">SCO</text><text x="114" y="48" text-anchor="middle">COU</text><text x="95" y="107" text-anchor="middle">UNC</text><text x="33" y="107" text-anchor="middle">VAL</text><text x="14" y="48" text-anchor="middle">DEL</text></g><text x="64" y="124" text-anchor="middle" font-family="sans-serif" font-size="8" fill="#0f172a">L · 90m</text></svg>
Merge will be attempted after this approval when explicitly authorized. ## Decision record (recomputable)
DECISION D-445-8801
TICKET ticket-445
HEAD_SHA 04a5494f267c5351cb6cfe26979f066ecce90b91
CORRELATION_ID local-semcod-koru-pr-704-ticket-445
ACTOR agent:ifuri-validator-agent[bot]
APPLIED_RULE P-CORE-015
INPUT author_login = "tom-sapletta-com"
INPUT observed_checks = ["standard packs / conformance=PASS","governance / enforce=PASS","governance / remote lifecycle=PASS","onedev/local-verify=PASS"]
INPUT required_checks = ["onedev/local-verify","standard packs / conformance"]
INPUT required_checks_source = "protected registry + GitHub applied rules (env/request)"
INPUT reviewer_login = "ifuri-validator-agent[bot]"
INPUT semantic_review_assessment = {"schema":"subactor.validator/semantic-review-assessment/v1","subject":{"repository":"semcod/koru","pull_request":704,"head_sha":"04a5494f267c5351cb6cfe26979f066ecce90b91","base_sha":"552ed81026c9213d05029c9ce395d2e01fa0ca7a","diff_sha256":"fa532f63a2602f11d7f4add2255e7652611df9abb6f00db9bb838c385927a424"},"policy":{"policy_schema":"subactor.validator/semantic-review-policy/v1","policy_version":1,"policy_sha256":"676cb4516bbfed2a000e40b9b1b6e4a430ecc761ec546aeb53d721a1905cfdd7","required":false,"critical_paths":[],"observed_paths":["project/TICKETS.md","project/ticket-445/README.md","project/ticket-445/intent.json","src/koru/queue/runner.py","tests/test_queue_model_provenance.py"]},"grounding":"full-diff-not-per-finding-proof","execution_authority":false,"status":"not_required","reason":null,"review_sha256":null,"unresolved":[]}
INPUT superseded_checks = ["smoke"]
INPUT ticket_radar_receipt = {"schema":"subactor.ticket-radar/v1","base_sha":"552ed81026c9213d05029c9ce395d2e01fa0ca7a","head_sha":"04a5494f267c5351cb6cfe26979f066ecce90b91","change_digest":"35847f43bdfdc0a5d78a18c59a722822b77a29089d4d4f5653aadec63fceb432","score":68,"complexity":"L","estimated_minutes":90,"split_recommended":true,"services":[],"authority":"ADVISORY","promotion":"FORBIDDEN"}
VERDICT APPROVE AUTHORITY DETERMINISTIC
REJECTED REQUEST_CHANGES BECAUSE NO_UNSAFE_CHANGE_REASON_FOUND
ADVISORY llm_verdict = "APPROVE" MODEL "openai/cursor-auto"
ASSERT VERDICT_AUTHORITY != "ADVISORY"

@ifuri-validator-agent
ifuri-validator-agent Bot merged commit e960b60 into main Oct 8, 2026
4 of 5 checks passed
@ifuri-validator-agent
ifuri-validator-agent Bot deleted the ticket/445-actual-runtime-model branch October 8, 2026 21:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant