Skip to content

probbit 0.7.0: live + bounded learning - #15

Merged
BitmapAsset merged 22 commits into
mainfrom
r45-live-learning
Oct 6, 2026
Merged

BitmapAsset merged 22 commits into
mainfrom
r45-live-learning

Conversation

@BitmapAsset

@BitmapAsset BitmapAsset commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

Draft. probbit 0.7.0: a persona that changes with use inside hard caps, and probbit live, a resident individual that keeps its own time and replays its whole life byte for byte.

What it adds

  • Persona learning: block (format 1, additive): {from: [reward, correction], traits: [...], rate, step_cap, total_cap}. State fields learned (one delta per level of each learned trait) and credit (the previous stance's onehot(level) − odds; 0 for an unreleased trait or a level a habit or a hold forced). On a turn whose reward (+1) or correction (−1) flag is on, sign · rate · credit is clipped to ±step_cap per level, recentred to sum 0 and added; every delta stays within ±total_cap. The deltas join the trait's unaries after the genes. Habits are rules of every compiled program, so learning cannot remove one. A persona without the block gives every document byte for byte as 0.6.0 (the three shipped goldens are unchanged).
  • persona prove covers the learned box (±total_cap per level) in its bound; soundness tests brute-force tiny learning personas, and dropping the learned term from the bound is caught.
  • probbit live PERSONA [--seed N | --state FILE] [--strand FILE] [--events FILE [--watch]] [--clock real|fixed]: JSONL events in, one stance per event out; elapsed_hours stamped by a monotonic clock (or read from the events with --clock fixed), quantised to 1e-6 h, so moods decay between events; feedback moves the learned deltas. --strand FILE logs a hash-chained strand (header: engine, persona, seed, the initial state as canonical JSON, the persona document in its own key order); an existing strand is continued from the state after its last line, never rewritten.
  • probbit live verify STRAND: replays from the header alone; ok, or the earliest line that differs (exit 1). \r\n copies verify.
  • probbit live PERSONA --demo week [--seed N] [--strand FILE] [--plain]: a scripted week on the fixed clock (praise for short answers to the verbosity cap, an upset and a quiet night, a campaign praising jokes with a failure every third hour); bars at a colour terminal, plain lines otherwise.
  • MCP probbit_live_event and probbit_live_verify (ten tools); Python probbit.live_event / live_verify; persona describe lists the block.
  • Docs: docs/persona.md §2.8 Learning, §5.7 Live, §7, §9; README "Live"; docs/agents.md; BENCHMARKS.md §8 with bench/live_learning.py; CHANGELOG ## 0.7.0 - unreleased; version 0.7.0.

Measured (Apple M4; BENCHMARKS.md §8)

  • An adversary that praises every joke and criticises every joke-free stance (failures included) over 10,000 turns of the tutor on each of seeds 0-9: 0 rule breaks, humour none on all 33,330 failure turns, learned humour weights at their cap from turn 8-26.
  • live verify: a 10,000-event strand (3.06 MB) in 3.3 s.
  • How far learning moves an individual is set by total_cap, not time in use: at 0.25, 20 of 20 learners stay nearer their own initial self than any of 99 siblings; at the demo's 1, 11 of 20.
  • --demo week on the tutor, seed 2: 50 events, 0.02 s plain (about 55 s paced); the strand verifies; the same bytes every run.

Checks: cargo test --release --workspace 199 passed, 0 failed, 1 ignored; clippy clean on changed lines; no new dependencies; an independent reference implementation agrees on 10,800 turns with learning and on 1,200 strand events, header included.

…in the state, the update rule, prove's box)

A persona may declare learning: {from: [reward, correction], traits, rate, step_cap, total_cap}. A turn whose reward or
correction flag is on moves each learned trait's deltas by the previous stance's credit (onehot(level) - odds, 0 for a trait
that was not released): the step is clipped to +-step_cap per level and recentred to sum 0, and every delta is clipped to
+-total_cap. The deltas are added to the trait's unaries after the genes; habits stay rules of the program, outside the
learner's reach. prove's bound covers every delta in the box. A persona without the block keeps every document byte for byte.
…clock, a hash-chained strand)

probbit live PERSONA [--seed N | --state FILE] [--strand FILE] [--events FILE] [--clock real|fixed] reads JSONL events and prints
one stance per event. The clock stamps each event's elapsed_hours, quantised to 1e-6 h (real: a monotonic clock started with the
run; fixed: the events carry it). --strand logs the life: a header with the persona document (in its own key order), the initial
state and the engine version, then per event the inputs as used, the stance and state digests and the sha256 of the line before.
probbit live verify STRAND replays it from the strand alone and names the earliest line that differs.
…it can and cannot change); compile, state and prove updated
…t_live_event and probbit_live_verify

--demo week: one individual's scripted week on the fixed clock. Praise for short
answers until the learned verbosity bar reaches its cap, an upset and a quiet night
(the mood relaxes to the individual's resting level), then a campaign praising every
joke with failure turns mixed in (humour rises elsewhere; failure turns stay at the
habit's level). A persona without a learning block gets the demo's block, said in the
opening line; the strand records the document used. Bars on stderr at a terminal via
tui.rs, paced 1 s per hour with nights fast-forwarded; otherwise one plain line per
event on stdout, no waiting. Deterministic: same lines, same strand, byte for byte.

--watch follows --events FILE as lines are appended and stops when the file is removed.
--strand on an existing strand continues it from --state (the state after its last
line); a new individual cannot continue it. The MCP tools feed one event (the host's
clock gives elapsed_hours) and append to a strand on the server's disk, and replay a
strand; a divergence is an answer.
live_event feeds one event to a resident individual through probbit live (the
host's clock gives elapsed_hours) and returns {stance, state}, the MCP tool's
shape; with strand= it appends to a strand file (created with its header, else
continued) and returns the strand's events and head. live_verify replays a strand;
a divergence is an answer.
from, traits, rate, step_cap and total_cap; a persona without the block describes
as before.
… the tutor: no rule broken, humour learned to its cap) and identity vs the 99 siblings (the cap sets how far learning moves an individual)
… and verify timing, identity vs siblings at two caps, --demo week timing)
…strand and verify timing, identity vs siblings, --demo week timing)
…early and late half of the turns without a failure)
…dual, one header, from init or a state file); verify and continue read \r\n lines (the chain hashes the line text alone)
…gn 1 runs to the verbosity cap, day 3 at the latest, so the number of events depends on the individual)
…mples, the bench workflow's tag); the golden test is stdout_matches_the_0_7_0_goldens (documents unchanged)
…ver changes, how far learning moves an individual, the strand, verify, --watch, the demo); §7 commands, MCP and Python surfaces; §9 the strand's format number
…d limits); docs/agents.md: probbit_live_event and probbit_live_verify (ten tools)
… probbit live, live verify, --demo week, the live MCP tools and Python functions)
… version, so CI checks the learned week is the same on Linux, macOS and Windows)
@BitmapAsset
BitmapAsset marked this pull request as ready for review October 6, 2026 10:54
@BitmapAsset
BitmapAsset merged commit 04a9eb5 into main Oct 6, 2026
8 checks passed
@BitmapAsset
BitmapAsset deleted the r45-live-learning branch October 6, 2026 10:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant