Skip to content

Fail assert_unbiased on a deterministic estimator that misses the truth - #11

Merged
soodoku merged 2 commits into
mainfrom
claude/vibrant-heisenberg-osft1o
Sep 29, 2026
Merged

soodoku merged 2 commits into
mainfrom
claude/vibrant-heisenberg-osft1o

Conversation

@soodoku

@soodoku soodoku commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

assert_unbiased passed any study whose estimates never varied, however far they missed the truth, because zero Monte Carlo SE was reported as bias_t = 0. A single replicate off by 1e6 also passed. Now:

  • all-identical estimates equal to the truth → pass (unchanged)
  • all-identical estimates that miss the truth → ValueError: the study can't tell a deterministic bias from a discrete estimator that happened not to vary (per Codex review: a Bernoulli(0.001) estimator is all-zero in 400 draws ~2/3 of the time, so asserting "biased" would be wrong); bias_t is NaN
  • fewer than 2 replicates → ValueError, matching se_ratio_tolerance / assert_narrower
  • every study with nonzero spread → identical behaviour

Proof it fixes the bug

New tests in tests/test_negative.py (biased constant, rare-Bernoulli, single replicate) fail on main and pass here.

Proof of no regressions

  • Differential test: assert_unbiased from main and from this branch were run on the same 20,000 random studies (normal, Student-t, rare-Bernoulli, and constant estimators; varied reps and truths). 0 outcome changes among studies with spread. The only changes were 5,241 zero-spread studies that miss the truth, which went pass → raise, exactly the intended change.
  • Full suite: 88 passed, plus the src doctests; ruff, ruff format, pyright, pydoclint clean.
  • Downstream: streamcal's suite (which depends on simcheck) passes against this branch (144 passed).

🤖 Generated with Claude Code

https://claude.ai/code/session_011pXaJ5DLex7BrdUyfTqjL4

A study whose estimates never varied has a Monte Carlo standard error of
zero, and bias_t reported that as a t statistic of zero, so an estimator
returning 5.0 for a truth of 2.0 passed assert_unbiased. So did a single
replicate off by 1e6, because sampling_sd is zero below two replicates.

bias_t is now infinite for a constant that misses the truth and zero only
for one that hits it, and assert_unbiased refuses studies of fewer than
two replicates, as se_ratio_tolerance and assert_narrower already do.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011pXaJ5DLex7BrdUyfTqjL4
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-29T01:13:38.236948Z 1624af8 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9c254f4a63

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread src/simcheck/results.py Outdated
Zero observed spread does not imply a deterministic estimator: 400 draws
of a Bernoulli(0.001) estimator are all zero about two times in three,
and the previous commit reported that unbiased estimator as infinitely
biased. assert_unbiased now raises ValueError in that case, explaining
that the study cannot tell a deterministic bias from a discrete
estimator that did not vary, and bias_t is NaN rather than infinite.
A biased constant still cannot pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_011pXaJ5DLex7BrdUyfTqjL4
@soodoku
soodoku merged commit 5b27269 into main Sep 29, 2026
12 checks passed
@soodoku
soodoku deleted the claude/vibrant-heisenberg-osft1o branch September 29, 2026 02:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants