Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
98 changes: 96 additions & 2 deletions docs/rates-context-observer-research.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Ten-year rates context observation research

Status: `UNVALIDATED_RESEARCH`. This is one source-agnostic, standard-library pure observation function. It is not a downloader, enabled plugin, strategy policy, or production adoption. Existing modules, package exports, runner/catalog entries, wide external-context CSVs, defaults, dependencies, and runtime pins are unchanged.
Status: `UNVALIDATED_RESEARCH`. This module has source-agnostic, standard-library pure observation entry points. It is not a downloader, enabled plugin, strategy policy, or production adoption. Existing modules, package exports, runner/catalog entries, wide external-context CSVs, defaults, dependencies, and runtime pins are unchanged.

[中文合同](rates-context-observer-research.zh-CN.md)

Expand All @@ -10,7 +10,7 @@ Status: `UNVALIDATED_RESEARCH`. This is one source-agnostic, standard-library pu

The existing wide CSV merge coerces each column to numeric values. It cannot preserve this contract's source, revision, or publication metadata; do not insert these records into that merge and assume provenance survives. This module neither extends `DEFAULT_FRED_SERIES` nor modifies macro watch/actionable scoring.

## Input and configuration
## Existing v1 input and configuration

The input has exactly `schema_version=qsl.rates-context-input.research.v1`, `decision_at`, and `series`. `decision_at` is the actual evaluation/receipt decision time declared by the caller, with an explicit ISO timezone. `series` contains only the optional roles `nominal_10y`, `real_10y`, and `breakeven_10y`; a missing role yields explicit unknown. All three roles are ten-year measurements, not ETF prices or another tenor relabeled as ten-year.

Expand Down Expand Up @@ -72,3 +72,97 @@ Focused tests can run with `python -m unittest discover -s tests -p test_rates_c
Synthetic coverage includes percent-to-bp arithmetic, independent breakeven, individual source delays, negative yields, future-row/revision invariance, missing publication/receipt evidence, today's import of old data, nonfinite/Boolean/string values, wrong units, mixed sources/methodologies, reversed dates, visible duplicate revisions, wrong/missing endpoints, stale observations, invalid configuration, strict JSON, input immutability, and absent trading fields.

This is a preparation step. A future authorized collector/consumer integration must retain the row-level evidence rather than using the lossy wide CSV, then validate genuine source coverage/availability, strategy-specific economic usefulness, and approved consumption. Historical breadth membership/prices and NDX participation are separate gaps; this phase implements neither breadth nor an NDX proxy. No runtime adoption follows from a pure function or passing synthetic tests.

## Explicit forward-known v2 entry point

build_rates_context_observation_v2(snapshot, config) is a separate pure entry
point in the same module. It accepts only qsl.rates-context-input.research.v2
and returns qsl.rates-context-observation.research.v2. The original v1 entry
point, constants, input/output semantics and rejection behavior remain unchanged.
Neither entry point automatically upgrades the other's input. V2 never substitutes
first-seen for v1 available_at.

The v2 root has exactly schema_version, availability_basis, collector_id,
decision_at and series. availability_basis is collector_first_seen. collector_id
identifies the collector to which the declared first-seen time belongs: a nonempty,
whitespace-trimmed string of at most 256 characters. Existing source/series/basis,
percent-unit roles and explicit v1 window/age configuration are reused. There is
no MSS family, catalog or runtime consumer registration.

Each v2 row has exactly:

- observation_date, value, revision_id
- source_published_at: explicitly null, never omitted or inferred
- first_seen_at: when the named collector first completely received this declared semantic row version
- received_at: when the research consumer received that fixed version
- capture_sha256: the declared hash of the original complete response retained when the row version was first seen
- row_sha256: the producer's declared immutable row-content reference

Both hashes require lowercase 64-hex syntax. The observer does not read the
response, recompute either hash, authenticate origin or clocks, establish an
earliest capture, or verify the producer's hashing algorithm. These are references
and consistency checks, not historical PIT proof. The collector retains original
bytes/capture records externally. Repeated downloads must not refresh an old
first_seen or replace its first-capture reference with the latest whole-file hash.
Corrections require separately retained versions; old accepted decisions are not
rewritten. This stateless function cannot enforce those rules across calls.

### Time projection and visible conflicts

V2 timestamps require YYYY-MM-DDTHH:MM:SS, optionally 1–6 fractional second
digits, followed by Z or ±HH:MM (offset hours <=23, minutes <=59). Greater
precision, comma fractions and second-bearing offsets are unsupported and remain
unknown; times are never truncated or rounded. Valid timestamps normalize to UTC.
Selected rows require
first_seen_at <= received_at <= decision_at; the UTC first-seen date cannot
precede the observation date. known_at is the later first-seen/receipt time,
which equals receipt under valid ordering. It is not a publication timestamp.

Dates outside the explicit window, future observation dates, and rows whose
first-seen or receipt is later than the decision are excluded before non-selector
fields. Hidden rows cannot affect counts, validation, identities, endpoint records
or numeric results. An unparseable selector cannot prove invisibility and produces
unknown unless another valid selector has already excluded the row.

Visible rows retain strict chronological order. Two rows for one date, including
exact repeats, are ambiguous: no sorting, deduplication, latest-wins or revision
selection. A duplicated endpoint has no selected endpoint metadata. Conflicting
content or timing under one observation/revision identity is unknown. Reuse of one
row hash for different observation dates or values within a series is unknown.
One capture hash may legitimately cover multiple different rows.

An old observation first captured today may be known today when its explicit
window/age policy permits it. It is never visible to an earlier decision. Age stays
observation-date age, not capture/receipt age; today's import does not refresh it.

### Output and assurance

V2 preserves collector/basis and each unambiguous endpoint's null publication,
first-seen, receipt, known-at, revision and hash references. There is no available_at
field. first_seen_delay_calendar_days compares date labels;
consumer_receipt_lag_seconds measures first-seen-to-consumer receipt. Neither is
source-publication latency or market-close-to-publication latency.

Missing/invalid fields, identities, times, nonfinite values, visible conflicts,
wrong roles/bases/units, stale observations or unavailable exact endpoints produce
unknown and no numeric change. Invalid configuration retains ContractError.
AI, opportunity and control fields are outside this exact input contract.

The same-source nominal/real approximate spread remains separate from independent
reported breakeven. Missing breakeven does not erase valid nominal/real declaration
facts or a valid pair difference; overall quality remains unknown. Every output
retains UNVALIDATED_RESEARCH, CALLER_DECLARATIONS_AND_CONSISTENCY_ONLY,
historical_pit_verified=false, backtest_eligible=false and
position_control_allowed=false.

This is offline contract preparation. Actual capture, source rights, retained
source bytes, genuine first-seen clocks, forward history, economic qualification
and approved consumption remain unverified. A later separately reviewed MSS
adapter can consume fixed v2 fields without changing v1. Real collection/adoption
need their own evidence; pre-collection historical availability is not recreated.

The existing focused unittest command runs original v1 plus synthetic v2 cases:
version/collector gates, arrival boundaries, old-date visibility without backdating,
stale age, hidden future/revision invariance, visible ambiguity, hash-reference
consistency, partial series, absent independent breakeven, strict JSON and unchanged
inputs. Pure tests do not prove external first-seen persistence.
80 changes: 78 additions & 2 deletions docs/rates-context-observer-research.zh-CN.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# 十年利率背景观察研究

状态:`UNVALIDATED_RESEARCH`。本期只有一个来源无关、标准库纯观察函数。不是下载器、已启用插件、策略政策或生产采用。既有模块、包级导出、runner/catalog、外部背景宽表CSV、默认配置、依赖和runtime pin保持原样。
状态:`UNVALIDATED_RESEARCH`。保留原来源无关、标准库纯 v1 观察入口,并新增独立显式 v2 入口。不是下载器、已启用插件、策略政策或生产采用。既有模块、包级导出、runner/catalog、外部背景宽表CSV、默认配置、依赖和runtime pin保持原样。

[Full English contract](rates-context-observer-research.md)

Expand All @@ -10,7 +10,7 @@

原宽表CSV合并会将各列转成数值,不能保存本合同的来源、修订和发布时间。不能把新记录塞进该合并后声称来源证据仍完整。本模块不扩充 `DEFAULT_FRED_SERIES`,不修改macro的watch/actionable评分。

## 输入与配置
## 原 v1 输入与配置

输入恰有 `schema_version=qsl.rates-context-input.research.v1`、`decision_at`、`series`。`decision_at` 是调用方声明的实际评估/接收决策时点,须有明确ISO时区。`series` 仅允许可缺省的 `nominal_10y`、`real_10y`、`breakeven_10y` 三个角色;缺少角色明确unknown。三者必须是真正十年期测量,不能把ETF价格或其他期限重标为十年。

Expand Down Expand Up @@ -72,3 +72,79 @@ focused测试可用 `python -m unittest discover -s tests -p test_rates_context_
合成覆盖percent→bp、独立breakeven、各源延迟、负利率、未来行/修订不变、缺公布/接收证据、今天导入旧数据、NaN/无穷/布尔/字符串、错误单位、混源/混口径、倒序/可见重复修订、错误/缺失端点、过时、非法配置、严格JSON、输入不变及无交易字段。

这只是准备步骤。未来获准采集/consumer接线须保留逐行证据,不能经过丢失metadata的宽表;另验真实来源覆盖/可得性、策略经济价值及获批消费。breadth历史membership/prices和NDX参与结构是独立缺口,本期没有实现breadth或NDX代理。纯函数或合成测试通过均不证明runtime采用。

## 显式 forward-known v2 入口

同模块新增独立纯入口 build_rates_context_observation_v2(snapshot, config),
仅接收 qsl.rates-context-input.research.v2,返回
qsl.rates-context-observation.research.v2。原 v1 入口、常量、输入/输出语义
及拒绝行为保持不变。两个入口不自动升级对方输入;v2 不把 first_seen
填入 v1 的 available_at。

v2 根字段恰为 schema_version、availability_basis、collector_id、decision_at、
series。availability_basis 固定 collector_first_seen。collector_id 是 first_seen
所归属采集器的声明身份,须为无首尾空白、非空且不超过 256 字符的字符串。
source/series/basis/percent 单位、三个角色和显式 v1 窗口/年龄配置继续沿用。
本批不注册 MSS family、catalog 或运行消费者。

每个 v2 row 恰有:

- observation_date、value、revision_id
- source_published_at:必须显式 null,不能省略或推算
- first_seen_at:指定采集器首次完整收到该声明语义行版本的时间
- received_at:研究 consumer 收到该固定版本的时间
- capture_sha256:该行版本首次被看到时保留的原始完整响应之声明 hash
- row_sha256:producer 声明的不可变行内容引用

两个 hash 均要求小写 64 位十六进制形式。observer 不读取响应、不重算 hash、
不认证来源或时钟、不确定“首次”真实性,也不验证 producer 的 hash 算法。
这里是引用和一致性检查,不是历史 PIT 证据。collector 在外部保留原始 bytes
和捕获记录。重复下载不能刷新旧 first_seen 或把首次 capture 引用换成新版
整文件 hash;更正须另留版本,不改写旧接受决策。无状态纯函数不能跨调用
强制这些持久记录规则。

### 时间投影与可见冲突

v2 时间精确支持 YYYY-MM-DDTHH:MM:SS,可带 1–6 位小数秒,后接 Z 或 ±HH:MM。
偏移小时不得大于 23、分钟不得大于 59;不支持更高精度、逗号小数或带秒的偏移。
不截断/舍入时间;不支持的形式保持 unknown。合法时间归一 UTC。选中行满足 first_seen_at <= received_at <= decision_at;
first_seen 的 UTC 日期不得早于观察日。known_at 为 first_seen/received 较晚值,
合法顺序下等于 received,不是源发布时间。

先排除显式窗口外日期、未来观察日,以及 first_seen 或 received 晚于决策的行,
再处理非选择器字段。隐藏行不影响计数、校验、身份、端点记录和数值。
无法解析的选择器不能证明不可见;若其他合法选择器尚未排除该行,则 unknown。

可见行保持严格日期顺序;同日两行,包括完全重复,也属于歧义。
不排序、不去重、不 latest-wins、不自行选择修订。重复端点没有选中端点 metadata。
同一观察/revision 身份下内容或时间冲突为 unknown;同一 series 中,
同一 row hash 对应不同观察日或数值也为 unknown。一个 capture hash 可以
合法对应同份响应的多条不同记录。

旧观察今天首次采集后,可以在显式窗口/年龄允许时成为今天已知的信息,
不能成为更早决策可见的信息。年龄始终取观察日期,不按采集/接收时间刷新。

### 输出与保证范围

v2 保留 collector/basis 及每个无歧义端点的 null publication、first_seen、
received、known_at、revision 和 hash 引用,不输出 available_at。
first_seen_delay_calendar_days 仅比较日期标签;consumer_receipt_lag_seconds
表示首次采集到 consumer 接收的间隔。两者都不是源公布延迟或市场收盘至发布时间。

缺失/非法字段、身份、时间、非有限值、可见冲突、错误角色/basis/单位、
过期或准确端点不可用,均 unknown 且不给数值变化。非法配置仍为 ContractError。
AI、机会和控制字段不属于该精确合同。

同源 nominal/real 的 approximate spread 与独立 reported breakeven 分开。
缺 breakeven 不抹掉两条有效声明事实或合法配对差值,整体 quality 仍 unknown。
所有输出保持 UNVALIDATED_RESEARCH、CALLER_DECLARATIONS_AND_CONSISTENCY_ONLY、
historical_pit_verified=false、backtest_eligible=false、position_control_allowed=false。

本批仅离线合同准备,未证明真实采集、来源权利、原始 bytes、可信 first_seen
时钟、前向历史、经济资格或获批消费。冻结 v2 后另行复审 MSS adapter 才接入,
不改旧 v1。真实采集/采用各需证据,不能还原采集前的历史可得性。

既有 focused unittest 命令同时运行原 v1 和合成 v2:版本/collector gate、
到达边界、旧日期非回填可见性、旧数据年龄、未来/迟到修订不变性、
可见版本歧义、hash 引用一致性、局部 series、缺独立 breakeven、严格 JSON
和输入不变。纯测试不能证明外部 first_seen 持久性。
Loading
Loading