A dependency-free Rust core with a separate JSON and SQLite agent adapter. It turns the techniques in the Superforecasting Skills Pack into composable APIs with explicit assumptions, validated inputs, and inspectable records.
This is a forecasting core, not a claim of elite predictive performance. That requires a prospective record on resolved questions against meaningful baselines.
Read online: Forecasting Handbook · Engineering Reference
Two detailed mdBooks ship with the source and local release:
- Forecasting Handbook: question design, outside views, evidence updates, decomposition, counterevidence, panels, aggregation, journals, scoring, calibration, validation, and decisions.
- Engineering Reference: architecture, numerical contracts, process integration, storage, recovery, performance, testing, operations, and executed examples for all 30 JSON commands.
CARGO_TARGET_DIR=target cargo build --workspace --locked --offline
python3 scripts/docs.pyThe rendered books are target/books/handbook/index.html and target/books/reference/index.html. The documentation gate executes Rust snippets and every JSON example, checks generated references against the runtime schema, and verifies rendered local links. Install mdBook 0.5.2 to build them; reading the release's included HTML needs only a browser.
use supercast::{Probability, bayes};
let prior = Probability::new(0.30)?;
let posterior = bayes::update(
prior,
Probability::new(0.80)?, // P(evidence | event)
Probability::new(0.20)?, // P(evidence | no event)
)?;
assert!((posterior.get() - 0.631578947368421).abs() < 1e-12);
# Ok::<(), supercast::Error>(())Run locally:
cargo test --workspace
cargo clippy --workspace --all-targets -- -D warnings
cargo run --example forecast
cargo bench --bench core
cargo doc --workspace --no-depsIf your environment overrides Cargo's target directory to a read-only location, prefix commands with CARGO_TARGET_DIR=target.
The supercast core has no runtime or development dependencies. Add it to a consuming project with supercast = { path = "/path/to/supercast" }. The separate supercast-agent crate uses Serde, Schemars, and bundled SQLite; the workspace lockfile pins its dependencies. Build offline once dependencies are cached. Neither crate has been published to a registry.
CARGO_TARGET_DIR=target cargo build -p supercast-agent
python3 examples/agent_demo.pyThe demo runs a synthetic forecast through a temporary database, restarts the process, and checks that retrying a saved update returns the original receipt. To retain the example journal:
target/debug/supercast-agent --db forecasts.sqlite3 < examples/agent-workflow.jsonl
target/debug/supercast-agent --schemaThe adapter exposes 30 operations over newline-delimited JSON, including reference classes, decomposition, decisions, panel summaries, forecast storage, scoring, cluster-bootstrap uncertainty, and calibration selection with held-out evaluation. Writes are atomic and idempotent by request ID. Generated request schema, protocol guide, and Python subprocess example are included.
bash scripts/check.sh
python3 scripts/release.pyThe check runs formatting, all workspace tests, strict Clippy, API and mdBook checks, and the restart/retry demo. The release command repeats the checks, builds an optimized host binary, packages both offline books, a rebuildable source workspace, examples and dependency notices, runs the staged binary, and writes a checksummed archive under target/releases. Neither command publishes anything. Dependencies must be cached first (cargo fetch --locked on a connected machine). Python 3.11+ and mdBook 0.5.2 are required for the release script; the demo uses Python 3.8+.
See release scope and acceptance criteria for what code completion covers.
| Pack technique | Library primitives |
|---|---|
| Question design | workflow::Question: version, timeline, yes/no/void rules, source precedence |
| Reference classes | reference::ReferenceClass, empirical rates, bayes::BetaPrior, empirical quantiles |
| Decomposition | Conditional chains, scenario mixtures, intersection/union bounds, constant-hazard event model |
| Bayesian updates | Stable likelihood updates, LR sensitivity, origin-aware UpdateSession |
| Counterevidence | workflow::Challenge records contrary pathways and testable cruxes |
| Conditional trees | decompose::Indicator: two branches, coherence, entropy gain, sensitivities |
| Delphi elicitation | Human/model respondent metadata, missing responses, median and spread |
| Logit aggregation | Weighted logit and arithmetic pools, explicit alpha, recalibration mapping |
| Update ledger | Immutable question version, checked revision chain, cutoff validation, resolution history |
| Brier scoring | Binary/categorical Brier, log loss, ranked probability score, empirical CRPS |
| Calibration | Exact or binned diagnostics, raw/coarsened scores, binning residual, exclusion counts |
| Talent validation | Temporal/cluster holdout checks, matched comparisons, cluster-bootstrap intervals, time-grid scoring, calibration candidate selection |
| Integrated workflow | Runnable synthetic example connecting question, prior, evidence, updates, and resolution |
Decision utilities and expected value of information help an agent choose which research could change its action. These are additional primitives beyond the pack.
Probabilityis finite and in[0,1]. Logit pooling rejects exact endpoints; log loss preserves infinite penalties for confidently wrong outcomes.- Bayesian updates reject impossible evidence. Calculations use log space to avoid underflow from multiplying tiny likelihoods.
- Likelihoods are conditional assumptions supplied by the caller. Source credibility is not automatically a likelihood ratio. Origin deduplication catches repeated IDs, not undisclosed dependence.
- Scenario weights must sum to one. Callers must establish that scenarios are disjoint and exhaustive. Conditional chains require conditional probabilities.
- Unresolved outcomes stay unresolved. Binary Brier is in
[0,1]; categorical Brier is in[0,2]; ranked score is unnormalized in[0,K-1]; CRPS has the outcome's units. - Calibration reports distinguish raw and coarsened Brier. Binning does not generally preserve the raw score decomposition.
- Question and evidence timestamps use UTC Unix seconds. Publication and observation dates must both respect the evidence cutoff. A timestamp supplied by a caller is not independently verified.
- Ledger entries are immutable through the public API. Corrections append a revision; a changed target needs a new question version. Final yes/no/void resolutions are terminal; an unresolved adjudication may be followed by a later resolution.
- Pooling alpha and recalibration coefficients are explicit inputs. Fit them on separate data, then evaluate once on an untouched holdout. The library never declares a forecaster certified.
The GJP replay and synthetic stress harness tests real historical forecast aggregation against independent formulas and exercises JSON/SQLite behavior at increasing workloads. See observed benchmark results for cohort definitions, measured results, and limits.
Scalar updates and scoring allocate no memory. Pooling takes O(n) time and O(1) extra space. Streaming Brier stores only a count and running mean. Time-grid scoring is O(updates + grid points). Empirical CRPS sorts the supplied slice in O(n log n), avoiding the quadratic pairwise distance matrix. Calibration uses a sparse map; ledger ID lookup uses a hash set. The SQLite adapter incrementally validates consecutive writes using one connection-local journal cache, invalidated by external commits. Reads and cold mutations retain full replay; see storage profiling and results.
cargo bench --bench core measures the local implementation with warmup and optimizer barriers. It is a small timing harness, not a statistical benchmark suite or a cross-language speed comparison. See benchmark results.
The library supplies typed data and calculations; supercast-agent adds durable storage and JSON tools. An agent application supplies retrieval, source verification, model calls, scheduling, the factual adjudication decision, and its frozen evaluation protocol. Panel summaries do not conduct interviews, and challenge records do not perform counterevidence searches.
The core ledger is in-memory; the adapter persists and revalidates events in SQLite. Calibration training selects from a declared grid; it is not a general optimizer. Bootstrap intervals resample event clusters under caller-supplied independence assumptions. This release has no MCP server, native Python bindings, full Bayesian network, general distribution-fitting engine, or demonstrated prospective forecasting performance. The JSON protocol is application-specific and can be called from Python or another agent runtime through a subprocess.
See architecture and extension guide, source attribution, and the end-to-end example.