Fail assert_unbiased on a deterministic estimator that misses the truth - #11
Merged
Merged
Conversation
A study whose estimates never varied has a Monte Carlo standard error of zero, and bias_t reported that as a t statistic of zero, so an estimator returning 5.0 for a truth of 2.0 passed assert_unbiased. So did a single replicate off by 1e6, because sampling_sd is zero below two replicates. bias_t is now infinite for a constant that misses the truth and zero only for one that hits it, and assert_unbiased refuses studies of fewer than two replicates, as se_ratio_tolerance and assert_narrower already do. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011pXaJ5DLex7BrdUyfTqjL4
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9c254f4a63
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Zero observed spread does not imply a deterministic estimator: 400 draws of a Bernoulli(0.001) estimator are all zero about two times in three, and the previous commit reported that unbiased estimator as infinitely biased. assert_unbiased now raises ValueError in that case, explaining that the study cannot tell a deterministic bias from a discrete estimator that did not vary, and bias_t is NaN rather than infinite. A biased constant still cannot pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_011pXaJ5DLex7BrdUyfTqjL4
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
assert_unbiasedpassed any study whose estimates never varied, however far they missed the truth, because zero Monte Carlo SE was reported asbias_t = 0. A single replicate off by 1e6 also passed. Now:ValueError: the study can't tell a deterministic bias from a discrete estimator that happened not to vary (per Codex review: a Bernoulli(0.001) estimator is all-zero in 400 draws ~2/3 of the time, so asserting "biased" would be wrong);bias_tis NaNValueError, matchingse_ratio_tolerance/assert_narrowerProof it fixes the bug
New tests in
tests/test_negative.py(biased constant, rare-Bernoulli, single replicate) fail onmainand pass here.Proof of no regressions
assert_unbiasedfrommainand from this branch were run on the same 20,000 random studies (normal, Student-t, rare-Bernoulli, and constant estimators; varied reps and truths). 0 outcome changes among studies with spread. The only changes were 5,241 zero-spread studies that miss the truth, which went pass → raise, exactly the intended change.srcdoctests;ruff,ruff format,pyright,pydoclintclean.🤖 Generated with Claude Code
https://claude.ai/code/session_011pXaJ5DLex7BrdUyfTqjL4