feat(sleep): add adversarial candidate probes - #263
Bogdan (Dan) Baciu (bogdanbaciu21) wants to merge 4 commits into
Conversation
|
Ready for review. This is one commit and nine files on the current base, at exact head The focused integration slice passed on Ubuntu, macOS, and Windows: 177 passed, 2 optional skips, and 6 subtests on each runner. The complete suites also passed with zero failures: Ubuntu and macOS 1,513 passed; Windows 1,468 passed. Strict docs passed on Python 3.12. The CLA check passed. The upstream CI run is marked |
|
Thanks for the unusually thorough test matrix and for making the feature opt-in. The direction is useful, but the blocking decision is not yet sound enough to gate adoption. The main issue is attribution. Please make blocking baseline-relative: evaluate baseline and candidate on identical source/probe pairs, then base rejection on a paired change such as the candidate gap worsening relative to the baseline gap. Evidence should retain all four scores so the decision is auditable. At minimum, add regressions where (1) baseline and candidate are equally frame-sensitive and adoption is not blocked, (2) the candidate improves both scores but retains a gap and is not blocked, and (3) only candidate-introduced degradation is blocked. There are two related validity problems:
Until baseline-relative comparison, stochastic robustness, and conservative transformations are covered, please remove or disable the blocking path and keep the probes advisory-only. The deterministic mock tests and cross-platform green suite verify plumbing, but they do not validate the blocking signal itself. |
|
Got it. Will do. These are incredibly helpful comments thank you so much for the time and care you took to provide them. It will take me a bit to make these changes, will reply when complete at a high quality level. |
|
Thanks again for the review. All three issues are addressed at head
Validation at the exact head across five native runners (Ubuntu 3.10/3.11/3.12, macOS arm64, Windows): full suite 1,538 passed on Ubuntu and macOS and 1,493 on Windows with zero failures; the focused slice is 202 passed everywhere; strict docs pass. On your closing point: we considered removing the blocking path entirely and keeping probes advisory-only. We kept blocking opt-in behind the three conditions you set (baseline-relative comparison, repeated-rollout consistency, conservative transformations), plus the rollout floor, and advisory remains the default. If you would still prefer advisory-only until the margin has been calibrated on a public scenario, we are happy to disable blocking in this PR and propose it separately with that calibration. |
|
Thanks for fixing baseline-relative comparison, repeated rollouts, and conservative request transformations. Re-reviewing
I reproduced this using the real cycle/group pipeline and the PR's deterministic robust backend, with: The ordinary consolidation runs, but the skill group is reported as Please forward the rollout setting through the grouped call and per-group diagnostic config, and add a real multi-skill cycle regression with blocking enabled and at least two rollouts. This can be fully offline; no new paid-model experiment is necessary for this wiring fix. |
…ation - blocking is baseline-relative: identical source/probe pairs scored under baseline and candidate docs; a row is brittle only when the candidate gap worsens beyond the margin in a strict majority of rollout indices - evidence retains all four aggregated scores plus per-rollout samples - dream_adversarial_rollouts (cap 8); blocking requires >= 2 - _strip_polite_frame restricted to politeness-marked requests with negative tests for ability/permission/desire forms - baseline documents are required arguments so no caller can silently compare against an unintended baseline
84bbde9 to
67b2f16
Compare
|
Thanks for the precise reproduction. I fixed the multi-skill omission and added the real offline regression you requested at head
Proof at No paid-provider run was added because this is the fully offline fanout wiring defect from your review. Ready for re-review. |
|
Re-reviewed This is the right kind of evidence for the wiring defect; I am not requesting a paid-provider run merely to verify argument propagation. Keeping the configured rollout count in diagnostics is also useful. Please refresh the validation summary in the PR description: its multi-platform table still says every job asserted The stated limitations on request-frame probes, proxy coverage, rollout cost, and opt-in blocking remain important. The focused propagation blocker is resolved; final review and the exact-head upstream CI gate remain. That CI run currently awaits maintainer approval, so this comment is not merge approval. |
|
Thanks. The validation summary is refreshed; the head is unchanged at
The stated limitations on request-frame probes, proxy coverage, rollout cost, and opt-in blocking stand as written. Nothing else changed. |
|
Thank you for addressing the multi-skill rollout handling and documenting the limits of the probe evidence. Re-reviewed One output mismatch remains at An independent real-cycle regression produces correct structured evidence: Its Please update the table columns and field names to expose the baseline-relative evidence, including |
|
Welcome back! Will do! |
|
Fixed the report mismatch in The table now shows baseline source/probe, candidate source/probe, and gap change. Four real-cycle regression cases cover brittle and robust candidates in advisory and blocking mode, asserting both the structured evidence and the exact numeric row in The reproduced brittle row now reads: Validation on that exact head: 51 adversarial tests pass locally; the five fork jobs (Ubuntu Python 3.10/3.11/3.12, macOS arm64 3.12, Windows 3.12) pass the focused integration slice and full suite. The full suite reports 1,547 passed on Ubuntu/macOS and 1,502 passed on Windows, with platform/optional skips reported separately. Strict docs pass on all three Python 3.12 platforms. The PR description now separates this receipt from the historical matrix. These are fork-runner results; upstream CI for this head still awaits repository approval. No paid-provider experiment was added for this rendering correction. This rendering fix is ready for re-review. Upstream workflow approval and the final merge decision remain with the project. |
What Problem This Solves
A candidate skill can improve its held-out validation score while depending on the exact wording or framing of harvested requests. It can therefore pass the ordinary gate and still break when the same request arrives with a harmless surface change.
The first revision of these probes scored only the candidate, so blocking mode could not distinguish brittleness the candidate introduced from prompt-frame sensitivity already present in the backend, and a single stochastic sample could mark a row brittle at the default margin.
Why This Change Was Made
This revision compares probe sensitivity with the baseline and adds a repeated-rollout consistency screen before any probe result can block adoption.
Identical source and probe pairs are scored under both the current documents and the candidate documents, each score is the mean of a configured number of repeated rollouts, and a row is brittle only when the candidate's probe minus source gap worsens beyond the margin relative to the baseline gap and the worsening holds in a strict majority of rollout indices. Blocking additionally requires at least two rollouts, so one stochastic sample can never reject a candidate; advisory runs may use one. The baseline documents are required arguments, so no caller can silently compare against an unintended baseline. Request-frame transformations are restricted to a defensible semantic-preservation contract: only explicitly politeness-marked requests are reframed, and ability, permission, and desire questions are never touched.
Project Fit
User Impact
Operators can see which harmless request variation broke a staged candidate, and whether the breakage is candidate-introduced or pre-existing backend sensitivity. Advisory mode adds evidence without changing gate decisions; explicit blocking mode rejects only candidate-introduced degradation, after operators calibrate
dream_adversarial_marginon their task mix and setdream_adversarial_rolloutsto at least two.Proof
The deterministic proof covers the pre-registered contract scenarios, each comparing the candidate arm versus the baseline arm on identical source and probe pairs.
Regression tests also pin train, validation, test, and provenance isolation; target-only routing in dual-backend operation; the three-variant, 256-probe, and eight-rollout resource bounds; strict configuration validation including the blocking rollout floor; non-finite and zero-probe fail-closed behavior; redaction; and Markdown-safe reporting.
These are deterministic contract tests, not a live-provider performance claim, so no stochastic lift or confidence interval is claimed.
Academic Support
Testing
Current head
4d93cad055f9210438867d02c8e7c24236902040The current head fixes the Markdown renderer to read all four baseline/candidate score fields and
gap_change. Four new real-cycle cases cover brittle and robust candidates in advisory and blocking mode. Each checks the structured scores and the exact numeric row in the generatedreport.md; all four failed on the prior head67b2f16441cfbefore the renderer fix.For the planted brittle candidate, the report now renders the baseline source/probe scores as
0.000 / 0.000, candidate source/probe as1.000 / 0.000, and gap change as-1.000, with statusbrittle.Fork-runner validation below asserts
actual_sha == expected_sha == 4d93cad055f9210438867d02c8e7c24236902040before testing. These are contributor-run checks, not official upstream CI.Local Linux/Python 3.12 validation also passed:
python3 -m pytest tests/test_adversarial_dream.py -qreported 51 passed;python3 -m pytest -qreported 1,547 passed, 12 skipped, 353 subtests; strict docs andgit diff --checkpassed. No test was deselected in the complete-suite runs and no failure was masked. The Windows count difference is reported as platform/optional-dependency skips.The exact-head upstream CI run is
action_requiredand awaits repository approval. No official upstream CI success is claimed.Historical evidence: superseded head 84bbde9, not current validation
The matrix below was collected on fork runners against
84bbde9ac191, the candidate reviewed before the multi-skill rollout-forwarding fix. It is retained for history only. It does not describe the current head and it is not official CI.Every job in that historical matrix asserted exact candidate
84bbde9ac191before testing. No test was deselected and no failure was masked. The Windows difference is explicit platform and optional dependency skips, not failures.Regression tests also pin train, validation, test, and provenance isolation; target-only routing in dual-backend operation; the three-variant, 256-probe, and eight-rollout resource bounds; strict configuration validation including the blocking rollout floor; non-finite and zero-probe fail-closed behavior; redaction; and Markdown-safe reporting.
Limitations & Negative Results
Reproduce It Yourself
Linux or macOS, shell in a fresh working directory: